Live23 Sept 2026/8 verified items today/Claim check: Yes, a 744B model runs on a laptop with no GPU/Daily digest
Preview: items marked "(demo)" are sample data, not yet editor-verified. How we verify

← Proofs · Claude Sonnet 5

Measured by us

RAG over documents — Claude Sonnet 5 on Direct API (global)

Disagree / report an error

Results

Answer accuracy89.8% higher is better
Citation faithfulness93.5% higher is better
Hallucination rate2.9% lower is better

Run details

Model version
claude-sonnet-5
Endpoint
api.anthropic.com /v1/messages, effort=high
Platform / region
Direct API / global
Dataset
docs-qa-500 (500 questions over 120 technical PDFs)
Harness
GitHub @ a1b2c3d
Run cost
$9.30
Date
23 Sept 2026

What was NOT tested

Re-run it on your data

git clone https://github.com/mallimatla/uptodate.git && cd uptodate && git checkout a1b2c3d
npm install && npm run proof -- --suite rag --model claude-sonnet-5 --platform direct --region global --dataset ./your-data

Sample raw output

The default is 3 retries with exponential backoff [doc 14, p. 7].