How we grade
We grade every tracker with letter grades, A to F, on each criterion and overall, so a tool ships with a report card like "Overall: A-". A criterion grade comes from a scored test run: we set each tool up on the same test brands, run our prompt panel against the assistants it claims to cover, and log coverage, data quality, and time to first useful report. The overall grade weights the criterion grades by the weights printed on the benchmark page.
Site views render scores on a 0 to 5 scale, so each letter maps to a fixed number: A is 5.0, A- is 4.7, B+ is 4.3, B is 4.0, B- is 3.7, C+ is 3.3, C is 3.0, D is 2.0, and F is 1.0. We also publish what we could not verify in each run. When a vendor gates a feature behind a sales call, we grade only what we could test ourselves.
Rerun cadence
We rerun the panel on a fixed schedule and after any material vendor change, such as new pricing or a new engine integration. The updated date on each page is the date the last run finished. Grades can move between runs, and the changelog on the benchmark page records every move with the reason.
Corrections and reruns
Every figure on the board is tied to a logged run and re-checked against the most recent completed run on a regular cadence. When a reading no longer matches, we rerun the affected test where the data warrants it and correct the page within two weeks.
See the latest ranking at /rankings/.