← Back to the public record explorer
benchmark event
Production AI public record - EditorialMethod →
GPT-4.1 — Q2 2026 Lab benchmark
GPT-4.1 scored 74/100 overall in the Q2 2026 PAI Lab PSF reliability index. Strong on structured output adherence. Notable gap: PII handling in summarisation tasks (PSF-03). Escalation trigger reliability above average.
Confidence
82%
Sources
2
Entities
2
Detected
Event summary
GPT-4.1 scored 74/100 overall in the Q2 2026 PAI Lab PSF reliability index. Strong on structured output adherence. Notable gap: PII handling in summarisation tasks (PSF-03). Escalation trigger reliability above average.
D1D2D3D4D6D5D7D8
Linked entities
Related graph edges
| Edge | Type | Confidence |
|---|---|---|
| ent-vendor-openai to ent-psf-d1 | maps to | 60% |
| ent-vendor-openai to ent-psf-d1 | maps to | 60% |
| ent-vendor-openai to ent-psf-d1 | maps to | 60% |
| ent-vendor-openai to ent-psf-d2 | maps to | 60% |
| ent-vendor-openai to ent-psf-d2 | maps to | 60% |
| ent-vendor-openai to ent-psf-d2 | maps to | 60% |
| ent-vendor-openai to ent-psf-d6 | maps to | 60% |
| ent-vendor-openai to ent-psf-d6 | maps to | 60% |
| ent-vendor-openai to ent-psf-d6 | maps to | 60% |
| ent-vendor-openai to ent-psf-d2 | maps to | 60% |
| ent-vendor-openai to ent-psf-d2 | maps to | 60% |
| ent-vendor-openai to ent-psf-d2 | maps to | 60% |