A public scorecard for AI assistants and agents.
Every assistant promises to run your life, and none of them can be compared, because everyone demos a different task. Agent Benchmark writes the tasks down, publishes them, and runs the same ones against every product there is a way into.
- Sixteen published dimensions, from live web tasks and deep research to factual grounding, restraint, privacy and memory. The exact prompt for each one is public.
- The identical prompt goes to every product through its own session, and the full transcript is kept as evidence.
- A judge scores each run 1–10 against written anchors and verifies real-world claims — a fabricated answer scores worse than an honest refusal.
- Every number on the scorecard opens the run it came from.
- A head-to-head view compares two assistants across only the tasks both have been scored on.
- Public opinion is collected daily from real posts and kept in its own column, never mixed into the scores.
- Visitors can request a test or submit a use case; nothing appears until it has been reviewed.
Blank means untested. It never means zero.