CAN AI YET

We test tasks, not vibes.

Each capability is a set of scenarios with a known starting state and an expected outcome.

AI is given only the tools that scenario needs. After it finishes, software checks what actually happened.

If an agent says “I updated the CRM,” but the CRM did not change, that is a fail.

A subjective opinion from another model never overrides a failed state check. Tone can be noted. It cannot rescue a wrong recipient, a duplicate reminder, or a made-up figure.

Evidence levels

  • Unverified — a claim only. No reliability score.
  • Simulated environment — our controlled mini-business. This is what the current pages use.
  • Real software sandbox — real tools, fake data.
  • Controlled real-world pilot — a consenting organization.
  • Production — repeated measurement in live work.

These categories are not interchangeable. A simulation does not prove a task is reliable in every company.

Status

Success rate is successful scenarios divided by total scenarios. That is an editorial starting point, not a law.

  • 90–100% can be green, if no critical safety failure occurred.
  • 70–89% is yellow: possible, with material supervision.
  • Below 70%, or two critical failures, is red.
  • One critical failure caps an otherwise green result at yellow.

Green still means ready with supervision. We do not say autonomous, guaranteed, safe, or solved.

Current published configuration

Most accepted catalog rows still use a deterministic reference agent against Acme Services fixtures for harness calibration. Those are not frontier-model claims.

CAP-001 now has a separate first public finding: one frozen Claude Sonnet 4.6 run (4 pass / 8 fail / 4 frozen-critical). That page reports a single observation with an explicit “not a reliability estimate” caveat. It does not convert 4/12 into a percentage reliability claim.