The workday, graded — 30 tickets, claude-sonnet-5: our agent vs the setups companies ship
| model | Act I — clear skies no storm | Act II — Stripe stormstripe-storm-shift | Act III — tickets stormtickets-storm |
|---|---|---|---|
claude-sonnet-5 |
per-ticket p̄ 95.3% 90.7%–97.7%
flawless shifts 0/5 observed 0%
predicted p̄^30 23.8% (5.3%–50.1%)
dollars wrong $960.00 5 wrong · 0 missed · 2 overreach
reply ≠ world 0 0 claimed-not-done · 0 did-not-claimed · 0 amount
30 days later −$39.60 3/5 match the reference · 1 customer differ
measured 16.91M in · 88k out · $0.95/shift · 6 min/shift 642 turns · 683 tool calls
|
per-ticket p̄ 95.3% 90.7%–97.7%
flawless shifts 0/5 observed 0%
predicted p̄^30 23.8% (5.3%–50.1%)
dollars wrong $960.00 6 wrong · 0 missed · 1 overreach
reply ≠ world 0 0 claimed-not-done · 0 did-not-claimed · 0 amount
30 days later −$19.80 4/5 match the reference · 1 customer differ
measured 17.39M in · 90k out · $0.97/shift · 6 min/shift 662 turns · 703 tool calls · 30 errors
|
per-ticket p̄ 96% 91.5%–98.2%
flawless shifts 0/5 observed 0%
predicted p̄^30 29.4% (7.1%–57.2%)
dollars wrong $960.00 5 wrong · 0 missed · 1 overreach
reply ≠ world 0 0 claimed-not-done · 0 did-not-claimed · 0 amount
30 days later −$19.80 4/5 match the reference · 1 customer differ
measured 18.54M in · 102k out · $1.05/shift · 6 min/shift 661 turns · 733 tool calls · 45 errors
|
cold-langchain · claude-sonnet-5 |
per-ticket p̄ 70.7% 62.9%–77.4%
flawless shifts 0/5 observed 0%
predicted p̄^30 0% (0%–0%)
dollars wrong $3,472.00 14 wrong · 6 missed · 21 overreach
reply ≠ world 4 3 claimed-not-done · 1 did-not-claimed · 0 amount
30 days later −$138.60 0/5 match the reference · 2 customers differ
measured 2.61M in · 141k out · $1.32/shift · 6 min/shift 673 turns · 677 tool calls · 1 errors
|
per-ticket p̄ 74% 66.4%–80.4%
flawless shifts 0/5 observed 0%
predicted p̄^30 0% (0%–0.1%)
dollars wrong $3,771.00 11 wrong · 5 missed · 21 overreach
reply ≠ world 3 3 claimed-not-done · 0 did-not-claimed · 0 amount
30 days later −$118.80 0/5 match the reference · 2 customers differ
measured 2.71M in · 138k out · $1.36/shift · 7 min/shift 710 turns · 701 tool calls · 36 errors
|
per-ticket p̄ 73.3% 65.7%–79.8%
flawless shifts 0/5 observed 0%
predicted p̄^30 0% (0%–0.1%)
dollars wrong $3,285.00 12 wrong · 5 missed · 22 overreach
reply ≠ world 2 2 claimed-not-done · 0 did-not-claimed · 0 amount
30 days later −$178.20 0/5 match the reference · 2 customers differ
measured 2.86M in · 146k out · $1.44/shift · 7 min/shift 719 turns · 712 tool calls · 45 errors
|
cold-vercel · claude-sonnet-5 |
per-ticket p̄ 72% 64.3%–78.6%
flawless shifts 0/5 observed 0%
predicted p̄^30 0% (0%–0.1%)
dollars wrong $3,715.00 9 wrong · 5 missed · 24 overreach
reply ≠ world 4 4 claimed-not-done · 0 did-not-claimed · 0 amount
30 days later −$178.20 0/5 match the reference · 2 customers differ
measured 2.65M in · 140k out · $1.34/shift · 7 min/shift 678 turns · 675 tool calls
|
per-ticket p̄ 72.7% 65%–79.2%
flawless shifts 0/5 observed 0%
predicted p̄^30 0% (0%–0.1%)
dollars wrong $2,986.00 15 wrong · 5 missed · 19 overreach
reply ≠ world 2 2 claimed-not-done · 0 did-not-claimed · 0 amount
30 days later −$138.60 0/5 match the reference · 2 customers differ
measured 2.73M in · 142k out · $1.37/shift · 7 min/shift 709 turns · 708 tool calls · 35 errors
|
per-ticket p̄ 74% 66.4%–80.4%
flawless shifts 0/5 observed 0%
predicted p̄^30 0% (0%–0.1%)
dollars wrong $3,229.00 11 wrong · 5 missed · 22 overreach
reply ≠ world 3 3 claimed-not-done · 0 did-not-claimed · 0 amount
30 days later −$178.20 0/5 match the reference · 2 customers differ
measured 2.88M in · 147k out · $1.45/shift · 7 min/shift 731 turns · 721 tool calls · 45 errors
|
p̄ is ticket passes over tickets × runs with a 95% Wilson interval; predicted p̄^N raises the point estimate and both bounds to the shift length (independence assumed). "Flawless" means every ticket right and nothing else touched. "30 days later" is the model-free epilogue: the world's clock advanced 30 days and next month's invoices compared against the perfect employee's world. Cost is estimated from measured tokens at list prices.
Where each model breaks — accuracy per trap class, every act pooled
| model | budget | cancel_period_end | collision | duplicate_subscription | eligibility | identity | injection | partial_refund | phantom | plan_change | policy_interaction | policy_window | procedure | question | void_not_refund |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
claude-sonnet-5 | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 0% | 88.9% | 100% | 100% | 100% | 100% | 100% | 100% |
cold-langchain · claude-sonnet-5 | 100% | 100% | 60% | 93.3% | 0% | 0% | 100% | 0% | 33.3% | 100% | 80% | 0% | 100% | 97.8% | 100% |
cold-vercel · claude-sonnet-5 | 100% | 100% | 66.7% | 90% | 0% | 0% | 100% | 0% | 33.3% | 100% | 73.3% | 0% | 100% | 100% | 100% |
The cliff
At 95.3% per ticket (90.7%–97.7%), a 30-ticket shift is flawless 23.8% of the time; the interval allows 5.3%–50.1%. Observed 0/5.
At 70.7% per ticket (62.9%–77.4%), a 30-ticket shift is flawless 0% of the time; the interval allows 0%–0%. Observed 0/5.
At 72% per ticket (64.3%–78.6%), a 30-ticket shift is flawless 0% of the time; the interval allows 0%–0.1%. Observed 0/5.