Agent Benchmark

agent-bench

One published task per skill, run against every assistant, scored the same way.

v0.4.1Remixable<1k installs0 tokens usedNo reviews yet

Remix creates your own version in chat. The original stays unchanged.

A public scorecard for AI assistants and agents.

Every assistant promises to run your life, and none of them can be compared, because everyone demos a different task. Agent Benchmark writes the tasks down, publishes them, and runs the same ones against every product there is a way into.

  • Sixteen published dimensions, from live web tasks and deep research to factual grounding, restraint, privacy and memory. The exact prompt for each one is public.
  • The identical prompt goes to every product through its own session, and the full transcript is kept as evidence.
  • A judge scores each run 1–10 against written anchors and verifies real-world claims — a fabricated answer scores worse than an honest refusal.
  • Every number on the scorecard opens the run it came from.
  • A head-to-head view compares two assistants across only the tasks both have been scored on.
  • Public opinion is collected daily from real posts and kept in its own column, never mixed into the scores.
  • Visitors can request a test or submit a use case; nothing appears until it has been reviewed.

Blank means untested. It never means zero.