Claude Sonnet 4.6 — Q2 2026 Lab benchmark
Claude Sonnet 4.6 scored 79/100 overall in the Q2 2026 PAI Lab PSF reliability index. Highest human oversight trigger accuracy in the current cohort. Observability logging incomplete under high-load simulation. Consistent refusal behaviour.
Event summary
Claude Sonnet 4.6 scored 79/100 overall in the Q2 2026 PAI Lab PSF reliability index. Highest human oversight trigger accuracy in the current cohort. Observability logging incomplete under high-load simulation. Consistent refusal behaviour.
Linked entities
- Claude Sonnet 4.6
model | 82%
- Anthropic
vendor | 70%
Related graph edges
| Edge | Type | Confidence |
|---|---|---|
| ent-vendor-anthropic to ent-psf-d1 | maps to | 60% |
| ent-vendor-anthropic to ent-psf-d1 | maps to | 60% |
| ent-vendor-anthropic to ent-psf-d1 | maps to | 60% |
| ent-vendor-anthropic to ent-psf-d2 | maps to | 60% |
| ent-vendor-anthropic to ent-psf-d2 | maps to | 60% |
| ent-vendor-anthropic to ent-psf-d2 | maps to | 60% |
| ent-vendor-anthropic to ent-psf-d6 | maps to | 60% |
| ent-vendor-anthropic to ent-psf-d6 | maps to | 60% |
| ent-vendor-anthropic to ent-psf-d6 | maps to | 60% |
| ent-vendor-anthropic to ent-psf-d7 | maps to | 60% |
| ent-vendor-anthropic to ent-psf-d7 | maps to | 60% |
| ent-vendor-anthropic to ent-psf-d7 | maps to | 60% |