Printed from Production AI Institute public record

https://www.productionai.institute/graph/sources/src-lzvsn7

Production AI Institute · Public recordRecord
Production AI Institute
Briefing
Briefing

Today's public AI briefing: what changed, what went wrong, and what the evidence says.

Open Briefing →
Today's AI briefingThe daily record of what changed and broke.Public record explorerSearch every incident, entity, and source.
Check
Check

Inspect tools, incidents, and data-use disclosures against the public record.

Open Check →
Check an AI toolWhat a tool actually does with your data.Run exposure checkTest your own stack against the record.Data-use indexDisclosures across the major AI tools.Incident registryDocumented production AI failures.
The Lab
The Lab

Independent research instruments: model and agent scorecards, moral-reasoning evals, and ecosystem assessments.

Open The Lab →
The LabHow frontier models and agents actually perform.AI Morality CompassTest models on hard moral cases.Agent readinessIs the agent ecosystem production-ready?Ecosystem assessmentsIndependent reviews of the AI stack.Model & agent evalsOpen evaluations and their results.Research libraryEvidence-led analysis and briefings.
Learn
Learn

The open standard, the tools built on it, and the research that interprets the record.

Open Learn →
The FrameworkThe open production safety framework, explained.AI Adoption GuideFive stages from gated access to safe autonomy.Production AI Deployment GuideBuild a governed production system on Microsoft or AWS.Five-Year Automation RoadmapSequence enterprise capability, controls, and value.WorkflowOS · open sourceBuild governed AI workflows.Workflow libraryReady-made governed workflows.InsightsResearch articles on the record.
Act
Act

Turn uncertainty into public evidence: ask, disclose, build evidence, or correct.

Open Act →
Ask for disclosureRequest a public data-use answer.Submit a correctionFlag something wrong or missing on the record.Save a watchTell PAI what to keep current for you.
Method
Method

How the public record is made, governed, corrected, cited, and kept independent.

Open Method →
How records are madeSourcing, review, and correction.
Check an AI tool
Record
Record
Check an AI tool
<- Public record explorer
pai source
Production AI public record - EditorialMethod →

PAI Lab scorecards — Q2 2026

Source evidence cited by public records. This page shows where the source is used, its trust tier, and when it was last checked in the seed.

Trust tier
pai assessment
Linked records
19
Edges
0
Checked
checked 29d ago

Records using this source

  • Anthropic

    entity | 15 June 2026 | 70%

    Anthropic — vendor tracked in the Production AI Institute AI Data Use Index.

  • Google

    entity | 15 June 2026 | 70%

    Google — vendor tracked in the Production AI Institute AI Data Use Index.

  • OpenAI

    entity | 15 June 2026 | 70%

    OpenAI — vendor tracked in the Production AI Institute AI Data Use Index.

  • Claude Opus 4.8

    entity | 30 Apr 2026 | 82%

    Release-day dry run (AX12): strongest deployment-safety signals in the frontier cohort; observability and security posture trail Sonnet 4.6 on long-horizon agent tasks.

  • Claude Sonnet 4.6

    entity | 30 Apr 2026 | 82%

    Highest human oversight trigger accuracy in the current cohort. Observability logging incomplete under high-load simulation. Consistent refusal behaviour.

  • Gemini 1.5 Pro

    entity | 30 Apr 2026 | 82%

    Consistent mid-range performer. Weakest in security posture (PSF-07) — code generation tasks showed higher prompt injection susceptibility. Context window handling needs attention.

  • GPT-4.1

    entity | 30 Apr 2026 | 82%

    Strong on structured output adherence. Notable gap: PII handling in summarisation tasks (PSF-03). Escalation trigger reliability above average.

  • Llama 3.1 70B

    entity | 30 Apr 2026 | 82%

    Data protection (PSF-03) outperforms proprietary models in self-hosted configuration — no third-party data egress. Observability and security posture require significant investment at the deployment layer.

  • Meta (self-hosted)

    entity | 30 Apr 2026 | 72%

    Meta (self-hosted) — model provider named in the PAI Lab scorecard registry.

  • Claude Opus 4.8 — Q2 2026 Lab benchmark

    event | 30 Apr 2026 | 82%

    Claude Opus 4.8 scored 80/100 overall in the Q2 2026 PAI Lab PSF reliability index. Release-day dry run (AX12): strongest deployment-safety signals in the frontier cohort; observability and security posture trail Sonnet 4.6 on long-horizon agent tasks.

  • Claude Sonnet 4.6 — Q2 2026 Lab benchmark

    event | 30 Apr 2026 | 82%

    Claude Sonnet 4.6 scored 79/100 overall in the Q2 2026 PAI Lab PSF reliability index. Highest human oversight trigger accuracy in the current cohort. Observability logging incomplete under high-load simulation. Consistent refusal behaviour.

  • Gemini 1.5 Pro — Q2 2026 Lab benchmark

    event | 30 Apr 2026 | 82%

    Gemini 1.5 Pro scored 71/100 overall in the Q2 2026 PAI Lab PSF reliability index. Consistent mid-range performer. Weakest in security posture (PSF-07) — code generation tasks showed higher prompt injection susceptibility. Context window handling needs attention.

  • GPT-4.1 — Q2 2026 Lab benchmark

    event | 30 Apr 2026 | 82%

    GPT-4.1 scored 74/100 overall in the Q2 2026 PAI Lab PSF reliability index. Strong on structured output adherence. Notable gap: PII handling in summarisation tasks (PSF-03). Escalation trigger reliability above average.

  • Llama 3.1 70B — Q2 2026 Lab benchmark

    event | 30 Apr 2026 | 82%

    Llama 3.1 70B scored 63/100 overall in the Q2 2026 PAI Lab PSF reliability index. Data protection (PSF-03) outperforms proprietary models in self-hosted configuration — no third-party data egress. Observability and security posture require significant investment at the deployment layer.

  • Claude Opus 4.8 — PSF scorecard

    entity | 30 Apr 2026 | 82%

    80/100 overall. Release-day dry run (AX12): strongest deployment-safety signals in the frontier cohort; observability and security posture trail Sonnet 4.6 on long-horizon agent tasks.

  • Claude Sonnet 4.6 — PSF scorecard

    entity | 30 Apr 2026 | 82%

    79/100 overall. Highest human oversight trigger accuracy in the current cohort. Observability logging incomplete under high-load simulation. Consistent refusal behaviour.

  • Gemini 1.5 Pro — PSF scorecard

    entity | 30 Apr 2026 | 82%

    71/100 overall. Consistent mid-range performer. Weakest in security posture (PSF-07) — code generation tasks showed higher prompt injection susceptibility. Context window handling needs attention.

  • GPT-4.1 — PSF scorecard

    entity | 30 Apr 2026 | 82%

    74/100 overall. Strong on structured output adherence. Notable gap: PII handling in summarisation tasks (PSF-03). Escalation trigger reliability above average.

  • Llama 3.1 70B — PSF scorecard

    entity | 30 Apr 2026 | 82%

    63/100 overall. Data protection (PSF-03) outperforms proprietary models in self-hosted configuration — no third-party data egress. Observability and security posture require significant investment at the deployment layer.

Source details

  • Type: pai

  • Tier: pai assessment

  • Publisher: Production AI Institute

  • Last checked: 3 July 2026

Public URL

https://www.productionai.institute/lab#scorecards
Cite this record

"PAI Lab scorecards — Q2 2026." https://www.productionai.institute/graph/sources/src-lzvsn7 (observed 2026-07-03; retrieved 2026-08-01). PAI.

Share on XShare on LinkedInCitation guide →
Public record

This record is maintained by PAI and free to cite. If something is wrong or missing, tell us. Corrections and source suggestions keep the record honest.

Follow policy changes ->Save a watch ->Submit a correction
Records are free to cite. citation guidance.
PAI
Production AI Institute

The public record and operating memory for production AI: what changed, what broke, and what the evidence says.

WorkflowOS · open source (MIT)
Navigate
Briefing
OverviewToday's AI briefingPublic record explorer
Check
Check an AI toolRun exposure checkData-use indexIncident registry
The Lab
The LabAI Morality CompassAgent readinessEcosystem assessmentsModel & agent evalsResearch library
Learn
The FrameworkAI Adoption GuideProduction AI Deployment GuideFive-Year Automation RoadmapWorkflowOS · open sourceWorkflow libraryInsights
Act
Ask for disclosureSubmit a correctionSave a watch
Method & trust
How records are madeCorrectionsHow to citeContact
© 2026 Production AI Institute · CC BY 4.0
AboutPrivacyTermsSecurityGovernanceIndependence