Position papers, incident analysis, framework assessments, and monthly intelligence that make production AI safety inspectable. The work is public because a standard only matters if people can inspect it.
Five developments shaping production AI this month - with the PAI angle on what they mean for practitioners building real systems today.
Read Issue 001 →Get the next issue in your inbox
Free. Monthly. Unsubscribe anytime.
Structured reliability testing of frontier AI models and agent frameworks against PSF criteria. Quarterly scorecards. Open methodology.
Explore →
Analysis, essays, and deep-dives on AI deployment, safety, and the production practitioner experience.
Explore →
Reusable workflow patterns for production AI systems - vetted against the PSF and ready to adapt.
Explore →
Documented AI failure cases with root-cause analysis mapped to PSF domains. Learn from what went wrong.
Explore →
Independent PSF assessments of every major AI framework - LangChain, CrewAI, AutoGen, Cursor, and more.
Explore →
The Production Safety Framework itself - eight domains, openly published and freely referenceable.
Explore →
Generate a PSF-aligned readiness report for an AI agent, with evidence grade, repository signals, and a shareable badge.
Explore →
May–June 2026 lab and ecosystem work - indexed here for procurement visitors; formal publication titles below are unchanged.
Lab-linked assessment of Cursor SDK 3.7 against PSF deployment-readiness and tool-permission controls.
Read →Forensic breakdown of the DN42 agent failure: control gaps, costs, and PSF-aligned prevention.
Read →How a tiny transfer exposed missing action gates, logging, and human oversight in a banking agent path.
Read →What broke, who was implicated, and which production controls the outage put under pressure.
Read →Anthropic confirmed elevated errors across Claude Opus 4.5 through 4.8 and Sonnet 4.6 on June 5, 2026. Verified timeline, PSF vendor-resilience controls, and wh
Read →Federal counsel used Anthropic Claude Console to draft a motion that included fabricated case quotations. Verified facts, PSF mapping, and controls for legal LL
Read →An independent PSF assessment of Anthropic Claude Opus 4.8 (May 28, 2026): effort control, dynamic workflows, and mid-task system entries. Strongest on agentic
Read →Cursor 3.6 (May 29, 2026) routes Shell, MCP, and Fetch calls through a classifier before they run. Our independent assessment: where the gate is real, where all
Read →Position papers, analyses, and framework notes from the PAI research programme.
Maps PSF domains to EU AI Act obligations for high-risk AI system deployers. Covers conformity assessment requirements, technical documentation standards, and human oversight obligations.
Analysis of common failure modes in production LLM deployments. Identifies root causes across PSF domains and intervention patterns.
Examines what constitutes meaningful human oversight in high-stakes AI-assisted decisions. Includes design patterns for effective human checkpoints.
Documents the reasoning behind each PSF domain, alternatives considered, and how practitioner feedback shaped the framework.
Patterns repeatedly observed in anonymised production assurance reviews and incident-led postmortems:
Model text consumed by downstream systems as if it were trusted structured data.
Review PSF-D2 →Strong model-call logs, weak cross-service traceability at queue and handoff boundaries.
See Lab methodology →Operational actions executed without explicit human intervention criteria on high-consequence paths.
Use deployment guidance →PAI collects anonymised incident reports from practitioners to inform framework development. If you have experienced a production AI failure and are willing to share details, we welcome your contribution.
Tool policy changes, new incident records, disclosure signals, and practical next steps. Public evidence, plain English, no hype.