Measured by usOutdated: newer version shipped, re-run pending
Agents and tool use — Claude Opus 5 on AWS Bedrock (us-east-1)
Results
| Task success rate | 83% higher is better |
| Median steps | 8 steps lower is better |
| Error recovery | 75.5% higher is better |
Run details
What was NOT tested
- Computer use
- Tasks longer than 50 steps
- Parallel sub-agents
Re-run it on your data
git clone https://github.com/mallimatla/uptodate.git && cd uptodate && git checkout a1b2c3d
npm install && npm run proof -- --suite agents --model claude-opus-5 --platform bedrock --region us-east-1 --dataset ./your-data