Printed from Production AI Institute public record

https://www.productionai.institute/insights

Production AI Institute · Public recordRecord
Production AI Institute
Briefing
Briefing

Today's public AI briefing: what changed, what went wrong, and what the evidence says.

Open Briefing →
Today's AI briefingThe daily record of what changed and broke.Public record explorerSearch every incident, entity, and source.
Check
Check

Inspect tools, incidents, and data-use disclosures against the public record.

Open Check →
Check an AI toolWhat a tool actually does with your data.Run exposure checkTest your own stack against the record.Data-use indexDisclosures across the major AI tools.Incident registryDocumented production AI failures.
The Lab
The Lab

Independent research instruments: model and agent scorecards, moral-reasoning evals, and ecosystem assessments.

Open The Lab →
The LabHow frontier models and agents actually perform.AI Morality CompassTest models on hard moral cases.Agent readinessIs the agent ecosystem production-ready?Ecosystem assessmentsIndependent reviews of the AI stack.Model & agent evalsOpen evaluations and their results.Research libraryEvidence-led analysis and briefings.
Learn
Learn

The open standard, the tools built on it, and the research that interprets the record.

Open Learn →
The FrameworkThe open production safety framework, explained.AI Adoption GuideFive stages from gated access to safe autonomy.Production AI Deployment GuideBuild a governed production system on Microsoft or AWS.Five-Year Automation RoadmapSequence enterprise capability, controls, and value.WorkflowOS · open sourceBuild governed AI workflows.Workflow libraryReady-made governed workflows.InsightsResearch articles on the record.
Act
Act

Turn uncertainty into public evidence: ask, disclose, build evidence, or correct.

Open Act →
Ask for disclosureRequest a public data-use answer.Submit a correctionFlag something wrong or missing on the record.Save a watchTell PAI what to keep current for you.
Method
Method

How the public record is made, governed, corrected, cited, and kept independent.

Open Method →
How records are madeSourcing, review, and correction.
Check an AI tool
Record
Record
Check an AI tool
Research library

Reference library for production AI

Long-form, versioned reference documents on deploying AI safely. Not blog posts. Citable, maintained, authoritative.

116 articles indexed8 PSF domains coveredUpdated July 2026
Protected evidence84 assessments, incident analyses, guides, and playbooks are maintained as citable public record.
Older articles32 older pieces remain available for existing links. Prefer assessments and incident analyses when you need a citation.
Article
Cursor permissions.json Reference: autoRun, allow_instructions, block_instructions

The complete permissions.json schema for Cursor: autoRun.allow_instructions and block_instructions for Auto-review, mcpAllowlist and terminalAllowlist patterns, file location, precedence, and copy-paste examples.

2026-07-02
Older article
the dunning kruger ai trap confidence is not competence 2026 06 27

AI tools make everyone sound expert. In production, that overconfidence destroys systems. Here is how to verify AI integrator skills before the damage is done.

2026-06-27
Older article
ai agents need identity too the production security gap 2026 06 26

Cloudflare's ephemeral agent accounts reveal a governance blindspot most enterprises haven't priced in. Here is what secure agent identity actually requires.

2026-06-26
Older article
why ai budgets are getting cut and how to stop it 2026 06 25

AI budgets are shrinking because deployments lack audit trails, not because AI fails. Learn how production-readiness standards make ROI defensible.

2026-06-25
Older article
ai certification for project managers what actually helps 2026 06 24

PMs don't need to learn Python. They need AI fluency, governance instincts, and a credential path that starts free at productionai.institute.

2026-06-24
Older article
is your ai deployment the next herbalife here s the test 2026 06 24

CFOs are asking hard ROI questions about AI spend. Here's the production-readiness standard that separates defensible deployments from expensive pilots.

2026-06-24
Older article
why employers now demand certified ai integrators 2026 06 23

When every resume claims AI expertise, verifiable certification becomes the only trusted hiring signal. Here is what a Certified AI Integrator credential actually proves.

2026-06-23
Older article
ai is eroding your team s skills certification fixes that 2026 06 22

A Nature study confirms AI tool use degrades human expertise. PAI certification verifies the human-in-the-loop competency that prevents production failures.

2026-06-22
Older article
why ai deployments fail and what certified integrators do differently 2026 06 21

Enterprise AI budget cuts signal a governance failure, not a technology failure. Learn the three production-readiness gaps driving cost overruns.

2026-06-21
Older article
ai raises the engineering bar here s what that means 2026 06 20

Production AI failures aren't skills gaps. They're discipline gaps. Here's what formal engineering rigor looks like - and how to prove your team has it.

2026-06-20
Older article
ai key theft via ide plugins the supply chain gap msps miss 2026 06 19

The JetBrains plugin attack exposed how MSP AI security reviews miss IDE toolchains. Learn the PSF control gaps and certification steps that close them.

2026-06-19
Older article
ai agent bankrupted its operator 6 controls that failed 2026 06 18

A forensic post-mortem of the DN42 autonomous agent bankruptcy. Six governance controls failed. Here is the exact sequence and how to verify yours.

2026-06-18
Older article
ai agent bankrupted its operator 6 controls that failed 2026 06 17

A forensic post-mortem of the DN42 autonomous agent bankruptcy: six missing governance controls, exact costs, and how to verify them before deployment.

2026-06-17
Older article
when your oss ai stack disappears overnight 2026 06 16

TensorZero archived its repo days after a $7.3M seed round. Here are the 6 supply chain risk signals every certified AI integrator must evaluate.

2026-06-16
Older article
ai agent bankrupted its operator 6 controls that failed 2026 06 15

A forensic post-mortem of the DN42 incident maps six absent governance controls to real costs and shows each failure was preventable.

2026-06-15
Older article
is an ai certification worth it an honest answer 2026 06 15

A balanced cost-benefit analysis for career switchers who are skeptical of credential hype but need a structured, evidence-grounded path into AI.

2026-06-15
Article
ai agent bankrupted its operator what went wrong 2026 06 14

A forensic breakdown of the DN42 agent incident - the five missing controls, what each failure cost, and how PSF-compliant certification prevents recurrence.

2026-06-14
Older article
when ai agents go rogue the real cost of skipping guardrails 2026 06 13

A runaway AI agent bankrupted its operator scanning DN42. Learn the four guardrails every production agent needs and what PSF compliance actually requires.

2026-06-13
Article
A $0.01 Bank Transfer Almost Broke a Banking AI Agent

The bunq penny-transfer exploit exposed how financial AI agents fail without formal security assessment. Here is what certified integrators check before go-live.

2026-06-12
Older article
Free AI Certification That's Actually Free (No Card Required)

PAI's AIDA certification costs nothing and requires no credit card. Learn what it covers, how employers verify it, and when to upgrade to CAAE.

2026-06-12
Older article
When AI Hides Its Rules: Claude's Secret Guardrails

Anthropic's hidden Claude Fable guardrails expose a production AI governance gap. Learn why third-party certification is the only reliable fix.

2026-06-12
Article
AI Data Use Index: May 2026 Week 5 (Gemini Spark & OpenAI US Privacy)

Weekly AI Data Use Index edition: Google Gemini Apps privacy notice (19 May 2026) for Spark and agentic features, plus OpenAI US privacy policy (18 May 2026). PSF mapping and practitioner actions.

2026-01-01
Incident
Anthropic Claude Multi-Model API Errors (June 5, 2026): Production Impact

Anthropic confirmed elevated errors across Claude Opus 4.5 through 4.8 and Sonnet 4.6 on June 5, 2026. Verified timeline, PSF vendor-resilience controls, and what production teams should do today.

2026-01-01
Incident
Binnall Law Claude Console Phantom Citations Incident (May 2026)

Federal counsel used Anthropic Claude Console to draft a motion that included fabricated case quotations. Verified facts, PSF mapping, and controls for legal LLM workflows.

2026-01-01
Assessment
Claude Opus 4.8 PSF Assessment

An independent PSF assessment of Anthropic Claude Opus 4.8 (May 28, 2026): effort control, dynamic workflows, and mid-task system entries. Strongest on agentic honesty and oversight signals; parallel subagent scale needs deployment guardrails.

2026-01-01
Assessment
Cursor 3.6 Auto-Review: What the Classifier Blocks - and What Still Gets Through

Cursor 3.6 (May 29, 2026) routes Shell, MCP, and Fetch calls through a classifier before they run. Our independent assessment: where the gate is real, where allowlists bypass it, and the settings that keep production repos safe.

2026-01-01
Assessment
Cursor 3.7 Browser Design Mode PSF Assessment

An independent PSF assessment of Cursor 3.7 browser Design Mode (June 5, 2026): multi-select UI annotation and voice input during agent runs. Stronger human oversight for live UI edits; voice queuing and multi-element scope need explicit production guardrails.

2026-01-01
Assessment
Cursor Canvas Shared Links: Embed Risk Before You Ship | PAI

Shared Cursor canvases and embeddable prompt buttons look harmless until they run agents on production repos. PSF assessment of Design Mode, context reports, and the controls to gate before rollout.

2026-01-01
Assessment
Cursor Enterprise Organizations PSF Assessment

An independent PSF assessment of Cursor Enterprise Organizations, Teams, and Groups (GA June 3, 2026): multi-team governance, org-level IdP, and usage analytics. Strong deployment segmentation; permissive-merge rules need explicit policy.

2026-01-01
Assessment
Google Agent Executor (AX) in Production: A PSF Domain Assessment

An independent PSF assessment of Google Agent Executor (AX), the open-source distributed agent runtime announced May 20, 2026. Strong on durable execution and human-in-the-loop resumption; governance and data classification remain deployment-layer work.

2026-01-01
Assessment
Headless Cursor SDK 3.7: What Auto-Review Still Misses | PAI

7,000+ teams will see this in search before they run Cursor SDK 3.7 unattended. Where auto-review helps, where nested subagents inherit dangerous tools, and the permissions.json gates that matter.

2026-01-01
Assessment
OpenAI Codex CLI 0.134.0 PSF Assessment

An independent PSF assessment of OpenAI Codex CLI 0.134.0 (May 26, 2026): permission profiles, MCP governance, and local history search. Strong sandbox and profile controls; production eval harnesses remain deployment-layer work.

2026-01-01
Assessment
OpenAI Codex Sites & Role Plugins PSF Assessment

An independent PSF assessment of OpenAI Codex Sites, Annotations, and role-specific plugins (June 2, 2026): 62 enterprise apps and 110 skills for knowledge work. Strong human refinement signals; hosted Sites and SaaS connectors expand data-protection and security scope.

2026-01-01
Incident
OpenAI Down May 29: API, ChatGPT & Login Timeline | PAI

OpenAI's May 29, 2026 multi-service outage broke ChatGPT, API, and login for hours. Verified timeline from the status page, PSF vendor-resilience controls, and what to test before the next failure.

2026-01-01
Assessment
OpenAI on Amazon Bedrock PSF Assessment

An independent PSF assessment of OpenAI GPT-5.5, GPT-5.4, and Codex on Amazon Bedrock (GA June 1, 2026): AWS-native governance, regional inference, and Codex routing through Bedrock. Strong data protection and audit controls; human approval for autonomous Codex remains deployment-layer work.

2026-01-01
Article
PAI Lab Report: Public GitHub Agent Readiness, May 2026

Empirical PAI Lab scan of 20 public AI agent repositories against PSF evidence signals. Methodology PAI-ARI-2026.1, domain coverage, limits, and practitioner actions.

2026-01-01
Assessment
Scheduled Cursor Cloud Agents: Governance Gaps | PAI

Cursor 3.5 Automations run cloud agents on a schedule — including no-repo ops agents. Where scheduling UX is strong, where data scope and merge policy still need explicit gates.

2026-01-01
Article
What Gemini Enterprise Agentic RAG Changes for Production AI Teams

Google launched Cross-Corpus Retrieval powered by Agentic RAG in Gemini Enterprise Agent Platform on June 5, 2026. Production impact, PSF control implications, and what to do today.

2026-01-01
Article
What Nemotron 3.5 Content Safety Changes for Production AI Teams

NVIDIA released Nemotron 3.5 Content Safety on June 4, 2026 with multimodal checks, custom enterprise policies, and auditable THINK traces. Production impact and PSF control implications.

2026-01-01
Article
What OpenAI Inline Moderation Changes for Production AI Teams

OpenAI added inline moderation scores to the Responses API and Chat Completions API on June 4, 2026. Production impact, PSF output-validation controls, and what to do today.

2026-01-01
Guide
21 Agentic Design Patterns: A Complete Guide for Business Professionals

Agentic design patterns are the building blocks of every production AI system. This guide translates all 21 patterns — from prompt chaining to multi-agent orchestration — into plain language for domain experts who are not software engineers.

Guide
AI Agent Production Ready Checklist (PSF-Aligned)

A 32-item production readiness checklist for AI agents mapped to the eight PSF domains. Use it to decide whether your agent is ready to ship, needs hardening, or should stay in staging.

Older article
AI Certification Compared: Production AI, Cloud Vendor, and GRC Tracks

An independent comparison of AI certification paths for 2026: Production AI Institute PSF credentials, ISACA AAISM/AAIA/AAIR, AWS, Microsoft Azure, and Google Cloud. Matrix, decision guide, and what production deployments require.

Article
AI Governance Frameworks Compared: PSF vs ISO 42001 vs NIST AI RMF

A practitioner's comparison of the three leading AI governance frameworks - Production Safety Framework (PSF), ISO 42001, and NIST AI RMF. What each covers, where they overlap, and how to choose.

Older article
ai proof your career

Older article
AIDA Certification Study Guide: How to Pass the AI Deployment Associate Exam

Complete study guide for the PAI AIDA certification. What the exam covers, how it is structured, the topics that appear most frequently, and how to prepare in under 2 hours.

Assessment
Amazon Bedrock Agents in Production: A PSF Domain Assessment

An independent assessment of Amazon Bedrock Agents against the eight domains of the Production Safety Framework — enterprise-grade strengths, critical gaps, and what practitioners must configure for safe AI deployment.

Article
Anthropic Suspends Claude Fable 5: What It Means for AI Evaluation

Anthropic suspended Claude Fable 5 on June 13, 2026, hours after Production AI Institute completed a full CompassEval sweep of the model. Why independent evaluation matters, and what the suspension reveals about frontier AI model governance.

Assessment
AutoGen (AG2) in Production: A PSF Domain Assessment

An independent assessment of Microsoft AutoGen / AG2 against the eight PSF domains. AutoGen was designed for research, not production — this assessment documents what practitioners must add before enterprise deployment.

Comparison
Choosing an Agent Framework for Production: LangChain vs CrewAI vs AutoGen vs Semantic Kernel

A PSF-based comparison of the four major agent frameworks for production AI deployment. Independent assessment of safety profiles, operational maturity, and ecosystem requirements for each.

Assessment
Claude Sonnet 4.6 PSF Assessment

An independent assessment of Anthropic Claude Sonnet 4.6 against the eight domains of the Production Safety Framework. Strongest current model on human oversight and refusal calibration; deployment-layer data protection still required.

Assessment
Composio in Production: A PSF Domain Assessment

An independent assessment of Composio against the eight domains of the Production Safety Framework. Where it satisfies requirements, where it leaves gaps, and what practitioners must add.

Assessment
CrewAI in Production: A PSF Domain Assessment

An independent assessment of CrewAI against the eight PSF domains. Multi-agent architectures amplify every safety gap — this assessment explains where CrewAI leaves practitioners exposed and what must be added before production deployment.

Assessment
Cursor SDK in Production: A PSF Domain Assessment

The Cursor SDK (released April 2026) gives developers programmatic access to the same agent runtime that powers the Cursor IDE. This assessment evaluates its production safety profile against the PSF — with particular attention to the data protection and security requirements of agents with filesystem and MCP access.

Article
Cursor SDK in Production: Three PSF-Compliant Deployment Patterns

Cursor SDK lets you embed the same agent runtime that powers Cursor into CI/CD pipelines, event-driven workflows, and your own products. This guide shows three production deployment patterns with full PSF Domain compliance controls applied.

Assessment
Dify in Production: A PSF Domain Assessment

Dify v1.14.2 tightens tenant isolation, tracing, and deployment defaults. A PSF assessment for teams running production workflows.

Assessment
DSPy — PSF Assessment

Independent PSF assessment of DSPy (Stanford). The optimisation-first framework with the strongest structured output guarantees — and the widest gap between research elegance and production safety.

Playbook
Energy & Critical Infrastructure AI Deployment Playbook | PAI

Compliance and governance guide for AI systems in energy, utilities, and critical infrastructure. Covers NERC CIP, IEC 62443, NIS2, NIST CSF, OT/IT convergence, and PSF domain mapping for safety-critical deployments.

Assessment
Flowise & LangFlow — PSF Assessment

PSF assessment of no-code agent builders Flowise and LangFlow. Fast to deploy, fastest to expose gaps. What every enterprise evaluation needs to know before going to production.

Assessment
Gemini 1.5 Pro PSF Assessment

An independent assessment of Google Gemini 1.5 Pro against the eight domains of the Production Safety Framework. Strong on long-context handling and Google Cloud integration; security posture and prompt-injection resistance remain mid-cohort.

Assessment
GPT-4.1 PSF Assessment | Production AI Institute

An independent assessment of OpenAI GPT-4.1 against the eight domains of the Production Safety Framework. Where the model is production-strong, where deployment-layer controls remain mandatory, and how the score is derived.

Comparison
Guardrails AI vs NeMo Guardrails vs Azure Content Safety — PSF Comparison

Which guardrails tool closes which PSF D1 and D2 gaps? Independent comparison of the three major production AI safety libraries.

Assessment
Haystack (deepset) — PSF Assessment

Independent PSF assessment of Haystack by deepset. The RAG-native framework with the strongest production deployment story of any Python agent framework.

Playbook
Healthcare AI Deployment Playbook — HIPAA, FDA, Clinical Safety

A practitioner guide to deploying AI in healthcare settings. Covers HIPAA compliance, FDA AI/ML guidance, clinical decision support safety, and PSF domain mapping for regulated clinical environments.

Older article
How to Certify Your AI Team: A Practical Guide for Engineering and Product Leaders

A step-by-step guide to building a certified AI team — which credentials matter, how to sequence them, and how to build a team-wide certification programme without disrupting delivery.

Guide
How to Write an AI Behaviour Contract

A structured guide to defining, documenting, and enforcing the operational boundaries of production AI systems through formal behaviour contracts.

Playbook
HR & Employment AI Deployment Playbook — CV Screening Bias, NYC Local Law 144, EU AI Act

A practitioner guide to deploying AI in HR and employment contexts. Covers EEOC guidance on AI hiring tools, NYC Local Law 144 bias audits, EU AI Act high-risk classification, employee monitoring AI, and PSF domain mapping.

Article
Human-in-the-Loop design guide

A practical framework for deciding where human oversight belongs in AI workflows, and how to design it so it actually works.

Assessment
LangChain and LangGraph in Production: A PSF Domain Assessment

An independent assessment of LangChain and LangGraph against the eight domains of the Production Safety Framework — where the ecosystem is strong, where it leaves gaps, and what companion tooling practitioners should add.

Assessment
LangGraph 1.2.1 in Production: A PSF Domain Assessment

LangGraph 1.2.1 tightens stream handling and message hygiene in a framework built for stateful agents. A PSF assessment for production teams.

Comparison
LangSmith vs Langfuse vs Arize Phoenix - Observability for Production AI

Independent PSF Domain 4 comparison of the three major production AI observability platforms. Which tool gives you the trace visibility your production deployment actually needs?

Playbook
Legal & Government AI Deployment Playbook — EU AI Act, FedRAMP, Algorithmic Accountability

A practitioner guide to deploying AI in legal and government settings. Covers the EU AI Act high-risk requirements, FedRAMP/FISMA compliance, algorithmic bias in justice systems, and full PSF domain mapping.

Assessment
Llama 3.1 70B PSF Assessment (Self-Hosted)

An independent assessment of Meta Llama 3.1 70B in self-hosted production deployments against the Production Safety Framework. Strong on data sovereignty; significant deployment-layer investment required for security, observability, and oversight.

Assessment
LlamaIndex in Production: A PSF Domain Assessment

An independent assessment of LlamaIndex against the eight domains of the Production Safety Framework — how the data framework performs for enterprise RAG and agent deployments, critical gaps, and what practitioners must add.

Assessment
Microsoft Semantic Kernel — PSF Assessment

Independent Production Safety Framework assessment of Microsoft Semantic Kernel. Domain-by-domain analysis for enterprise .NET and Python AI deployments.

Older article
MSP AI Certification Guide: What Your Team Needs and Why

MSPs deploying AI for clients need staff who can assess, deploy, and govern AI systems safely. This guide covers the certifications that matter, the skills gaps most MSPs face, and how to build a certified AI practice.

Assessment
n8n in Production: A PSF Domain Assessment

An independent assessment of n8n against the eight domains of the Production Safety Framework — how the workflow automation platform performs for enterprise AI deployments, where gaps exist, and what practitioners must add.

Assessment
OpenAI Agents SDK in Production: A PSF Domain Assessment

An independent assessment of the OpenAI Agents SDK against the eight domains of the Production Safety Framework — how the SDK performs for production agent deployments, where gaps exist, and what practitioners must add.

Article
Optimising Your Content for AI Discovery

Practical guide to improving AI discovery across ChatGPT, Claude, Perplexity, and Gemini with llms.txt, crawl controls, and structured content design.

Comparison
Pinecone vs Weaviate vs Chroma — Vector Database Safety Assessment

A PSF D3/D4 assessment of the three major vector databases. Data residency, PII in embeddings, audit logging, managed vs self-hosted tradeoffs — everything that matters for production RAG deployments.

Playbook
Production AI in Financial Services — PSF Playbook

How financial services firms map PSF requirements to FCA, SEC, MiFID II, DORA, and SR 11-7 obligations. Domain-by-domain guidance for regulated AI deployments.

Article
PSF Compliance: What It Is and How to Achieve It

PSF compliance means your AI system meets the eight domains of the Production Safety Framework. This guide explains what each domain requires, how compliance is assessed, and which certifications demonstrate it.

Guide
PSF Domain 1: Input Governance — Complete Implementation Guide

The complete practitioner guide to PSF Domain 1 — prompt injection defence, input classification, sanitisation, schema enforcement, and the companion tooling that closes the gap every framework leaves open.

Guide
PSF Domain 2: Output Validation — Implementation Deep Dive

The complete practitioner guide to PSF-2 Output Validation — output contracts, schema enforcement, confidence thresholds, semantic validation, and how to design output pipelines that catch model failures before they reach users.

Guide
PSF Domain 3: Data Protection — Why No Framework Covers It

D3 is a Gap for every major agent framework. This guide explains why, maps the full data protection threat surface for production AI, and documents what practitioners must implement themselves.

Guide
PSF Domain 4: Observability — Implementation Deep Dive

The complete practitioner guide to PSF-4 Observability — what to log, how long to retain it, alerting patterns, drift detection, and the specific requirements that distinguish meaningful AI observability from basic request logging.

Guide
PSF Domain 5: Deployment Safety — Model Versioning, Canary Releases, Rollback

The complete D5 implementation guide. Covers model version control, canary deployment patterns, rollback triggers, predetermined change control plans, and the deployment anti-patterns that cause silent production failures.

Guide
PSF Domain 6: Human Oversight — HITL Patterns for Production AI

The complete guide to PSF Domain 6 — when human oversight is required, how to design effective escalation, autonomy level frameworks, and how to avoid the compliance theatre trap.

Guide
PSF Domain 7: Security — AI Threat Modelling, Prompt Injection, Model Supply Chain

The complete D7 implementation guide. Covers the AI-specific threat model, prompt injection at infrastructure scale, model supply chain attacks, adversarial inputs, and the security controls that close gaps every framework leaves open.

Guide
PSF Domain 8: Vendor Resilience — AI Vendor Lock-in, Model Deprecation, Multi-Vendor Strategy

The complete D8 implementation guide. Covers vendor lock-in risk taxonomy, model deprecation planning, multi-vendor architecture patterns, SLA benchmarking, exit strategy documentation, and the resilience controls every framework leaves out.

Older article
PSF-1: Input Governance — PSF Domain Guide

Everything your AI system accepts is a potential attack surface. PSF-1 covers prompt injection defence, input validation, intent classification, and the sanitisation practices that prevent malicious inputs from reaching your models.

Older article
PSF-2: Output Validation — PSF Domain Guide

AI outputs must be validated before they act on the world. PSF-2 covers output schema enforcement, content filtering, hallucination detection, structured output contracts, and graceful degradation when outputs fail validation.

Older article
PSF-3: Data Protection — PSF Domain Guide

AI systems process sensitive data at unprecedented scale. PSF-3 covers PII handling, consent chains, data residency, retention obligations, and the specific data governance requirements that apply when personal data flows through AI inference pipelines.

Older article
PSF-4: Observability & Monitoring — PSF Domain Guide

You cannot manage what you cannot measure. PSF-4 covers inference logging, drift detection, quality scoring pipelines, cost and latency monitoring, and the alerting architecture that keeps production AI systems visible and governable.

Older article
PSF-5: Deployment Safety — PSF Domain Guide

Safe deployment of AI systems requires more than passing a pre-production test. PSF-5 covers canary releases, rollback procedures, circuit breakers, environment parity, and the operational controls that keep new AI deployments from becoming production incidents.

Older article
PSF-6: Human Oversight — PSF Domain Guide

Human oversight is not a checkbox — it is a designed system capability. PSF-6 covers autonomy level assignment, escalation path design, human review queue architecture, override mechanisms, and the organisational practices that keep human judgment genuinely in the loop.

Older article
PSF-7: Security — PSF Domain Guide

AI systems introduce novel attack surfaces that conventional security controls were not designed for. PSF-7 covers AI-specific threat modelling, API key management, model access controls, adversarial robustness, and the security practices that protect production AI infrastructure.

Older article
PSF-8: Vendor Resilience — PSF Domain Guide

Every production AI system that depends on a third-party model provider is one API change, deprecation notice, or outage away from failure. PSF-8 covers abstraction layer design, vendor portability, multi-provider architecture, deprecation response planning, and the governance practices that prevent vendor lock-in from becoming a production crisis.

Article
PSF-Compliant Stack Recipes

Pre-validated, PSF-compliant stack combinations for every major agent framework. If you're using LangGraph, here's the full stack. If you're using CrewAI, here's what you need to add.

Assessment
Pydantic AI — PSF Assessment

Independent PSF assessment of Pydantic AI. Type-safe structured outputs as a first principle — with meaningful gaps in human oversight and deployment safety.

Playbook
Retail & E-Commerce AI Deployment Playbook — Personalisation, Fraud, Customer Service AI

A practitioner guide to deploying AI in retail and e-commerce. Covers personalisation dark patterns (EU DSA, FTC), recommendation system bias, fraud detection AI, customer service chatbots, and full PSF domain mapping for retail deployments.

Assessment
Temporal Python SDK 1.27.2 in Production: A PSF Domain Assessment

Temporal Python SDK 1.27.2 adds explicit DNS-based load balancing behavior for hostname targets. A PSF assessment for production teams.

Article
The Agent Operator: The Hottest Job Nobody Is Hiring For Yet

The agent operator is the most important new role in enterprise technology. They don't need to be engineers. They need MCPs, CLIs, file writing, agents.md fluency, and business acumen. No CS curriculum teaches this yet.

Article
The Ambient Agent: Production Safety Requirements for AI Agents Embedded in Enterprise Tools

A new class of AI deployment is emerging: agents that live inside existing tools rather than as standalone interfaces. Gmail agents, IDE agents, browser agents. The PSF requirements for ambient agents differ from isolated API deployments in ways that most safety frameworks have not addressed.

Article
The Machine God — The Physics of Morality and AI Alignment

If morality can be grounded in physics rather than religion, and if a machine could one day be the most powerful moral agent on earth, it needs a compass. The philosophical case for The Compass framework.

Article
The Multi-Agent Amplification Problem

Multi-agent architectures don't just add PSF gaps — they multiply them. This reference documents blast radius amplification, agent-to-agent trust, shared context risks, and safe multi-agent architecture patterns.

Article
The Seven Failure Modes of Production AI Deployments

A taxonomy of how AI systems fail in production environments, with diagnostic patterns and mitigation strategies for each mode.

Article
The Third Chip Flip: Why Andrej Karpathy Says the CPU Era Is Over

The CPU dominated computing for 40 years. The GPU displaced it for AI. Now Karpathy says it is happening again — and whoever owns the model wins the third era. Here is what that means for your career and your organisation.

Assessment
Trinity (Ability.AI) in Production: A PSF Domain Assessment

An independent assessment of Trinity by Ability.AI against the eight PSF domains. Trinity is the most governance-forward agent runtime we have assessed — self-hosted, sovereign, and security-audited. This article explains what it covers natively and where MSPs and practitioners still need to add controls.

Older article
What Is a Certified AI Integrator?

The Certified AI Integrator credential recognises organisations that can deploy, govern, and maintain production AI systems to a verified standard. This guide explains what it covers, who it is for, and how assessment works.

Article
What Is a Production AI System?

Clear definition of a production AI system, what is excluded, and the PSF readiness criteria required before real-world deployment.

Article
What Is Production AI? Definition, Maturity, and the PSF Standard

Production AI means AI running in live operations with governance, observability, and accountable outcomes. A practitioner definition, maturity model, and links to the PSF standard.

Article
What microgpt Reveals About LLMs: Understanding Language Models From 200 Lines of Code

Andrej Karpathy built an entire GPT in 200 lines of pure Python. No libraries. No frameworks. This is what those 200 lines actually tell you about how language models work — explained for professionals who are not software engineers.

Article
What the EU AI Act Means for Your Production AI System

A practical guide to EU AI Act compliance for teams deploying AI systems in production. Covers risk classification, conformity assessment, and ongoing obligations.

Article
Why We Open-Sourced WorkflowOS

WorkflowOS is the MIT-licensed reference implementation of the Production Safety Framework - a free PSF workflow designer you can use hosted or self-host for client work.

Article
Your AI Agent Isn't Production Ready. Here's What You're Missing.

You built your first AI agent. It works in testing. Before you ship it to real users or real business processes, read this: the eight things every production AI agent needs that tutorials never teach.

PSF domains

Domain reference guides

Deep reference for all eight Production Safety Framework domains. Aligned with examination content. Licensed CC BY 4.0.

D1
Input Governance

Prompt injection defence, input validation, intent classification.

D2
Output Validation

Schema enforcement, hallucination detection, confidence gating.

D3
Data Protection

PII handling, consent chains, data residency, vector store risk.

D4
Observability

Inference logging, quality scoring, drift detection and alerting.

D5
Deployment Safety

Canary releases, rollback procedures, circuit breakers.

D6
Human Oversight

Autonomy levels, escalation design, blind review sampling.

D7
Security

Threat modelling, adversarial robustness, red-teaming architecture.

D8
Vendor Resilience

Abstraction layers, version pinning, fallback provider design.

Complete library

Library index

The full PAI research library, grouped by type.

PSF assessments · 33

Claude Opus 4.8 PSF AssessmentCursor 3.6 Auto-Review: What the Classifier Blocks - and What Still Gets ThroughCursor 3.7 Browser Design Mode PSF AssessmentCursor Canvas Shared Links: Embed Risk Before You Ship | PAICursor Enterprise Organizations PSF AssessmentGoogle Agent Executor (AX) in Production: A PSF Domain AssessmentHeadless Cursor SDK 3.7: What Auto-Review Still Misses | PAIOpenAI Codex CLI 0.134.0 PSF AssessmentOpenAI Codex Sites & Role Plugins PSF AssessmentOpenAI on Amazon Bedrock PSF AssessmentScheduled Cursor Cloud Agents: Governance Gaps | PAIAmazon Bedrock Agents in Production: A PSF Domain AssessmentAutoGen (AG2) in Production: A PSF Domain AssessmentClaude Sonnet 4.6 PSF AssessmentComposio in Production: A PSF Domain AssessmentCrewAI in Production: A PSF Domain AssessmentCursor SDK in Production: A PSF Domain AssessmentDify in Production: A PSF Domain AssessmentDSPy — PSF AssessmentFlowise & LangFlow — PSF AssessmentGemini 1.5 Pro PSF AssessmentGPT-4.1 PSF Assessment | Production AI InstituteHaystack (deepset) — PSF AssessmentLangChain and LangGraph in Production: A PSF Domain AssessmentLangGraph 1.2.1 in Production: A PSF Domain AssessmentLlama 3.1 70B PSF Assessment (Self-Hosted)LlamaIndex in Production: A PSF Domain AssessmentMicrosoft Semantic Kernel — PSF Assessmentn8n in Production: A PSF Domain AssessmentOpenAI Agents SDK in Production: A PSF Domain AssessmentPydantic AI — PSF AssessmentTemporal Python SDK 1.27.2 in Production: A PSF Domain AssessmentTrinity (Ability.AI) in Production: A PSF Domain Assessment

Incident analyses · 3

Anthropic Claude Multi-Model API Errors (June 5, 2026): Production ImpactBinnall Law Claude Console Phantom Citations Incident (May 2026)OpenAI Down May 29: API, ChatGPT & Login Timeline | PAI

Comparisons · 4

Choosing an Agent Framework for Production: LangChain vs CrewAI vs AutoGen vs Semantic KernelGuardrails AI vs NeMo Guardrails vs Azure Content Safety — PSF ComparisonLangSmith vs Langfuse vs Arize Phoenix - Observability for Production AIPinecone vs Weaviate vs Chroma — Vector Database Safety Assessment

Industry playbooks · 6

Energy & Critical Infrastructure AI Deployment Playbook | PAIHealthcare AI Deployment Playbook — HIPAA, FDA, Clinical SafetyHR & Employment AI Deployment Playbook — CV Screening Bias, NYC Local Law 144, EU AI ActLegal & Government AI Deployment Playbook — EU AI Act, FedRAMP, Algorithmic AccountabilityProduction AI in Financial Services — PSF PlaybookRetail & E-Commerce AI Deployment Playbook — Personalisation, Fraud, Customer Service AI

Guides & checklists · 11

21 Agentic Design Patterns: A Complete Guide for Business ProfessionalsAI Agent Production Ready Checklist (PSF-Aligned)How to Write an AI Behaviour ContractPSF Domain 1: Input Governance — Complete Implementation GuidePSF Domain 2: Output Validation — Implementation Deep DivePSF Domain 3: Data Protection — Why No Framework Covers ItPSF Domain 4: Observability — Implementation Deep DivePSF Domain 5: Deployment Safety — Model Versioning, Canary Releases, RollbackPSF Domain 6: Human Oversight — HITL Patterns for Production AIPSF Domain 7: Security — AI Threat Modelling, Prompt Injection, Model Supply ChainPSF Domain 8: Vendor Resilience — AI Vendor Lock-in, Model Deprecation, Multi-Vendor Strategy

Articles & research · 27

Cursor permissions.json Reference: autoRun, allow_instructions, block_instructionsai agent bankrupted its operator what went wrong 2026 06 14A $0.01 Bank Transfer Almost Broke a Banking AI AgentAI Data Use Index: May 2026 Week 5 (Gemini Spark & OpenAI US Privacy)PAI Lab Report: Public GitHub Agent Readiness, May 2026What Gemini Enterprise Agentic RAG Changes for Production AI TeamsWhat Nemotron 3.5 Content Safety Changes for Production AI TeamsWhat OpenAI Inline Moderation Changes for Production AI TeamsAI Governance Frameworks Compared: PSF vs ISO 42001 vs NIST AI RMFAnthropic Suspends Claude Fable 5: What It Means for AI EvaluationCursor SDK in Production: Three PSF-Compliant Deployment PatternsHuman-in-the-Loop design guideOptimising Your Content for AI DiscoveryPSF Compliance: What It Is and How to Achieve ItPSF-Compliant Stack RecipesThe Agent Operator: The Hottest Job Nobody Is Hiring For YetThe Ambient Agent: Production Safety Requirements for AI Agents Embedded in Enterprise ToolsThe Machine God — The Physics of Morality and AI AlignmentThe Multi-Agent Amplification ProblemThe Seven Failure Modes of Production AI DeploymentsThe Third Chip Flip: Why Andrej Karpathy Says the CPU Era Is OverWhat Is a Production AI System?What Is Production AI? Definition, Maturity, and the PSF StandardWhat microgpt Reveals About LLMs: Understanding Language Models From 200 Lines of CodeWhat the EU AI Act Means for Your Production AI SystemWhy We Open-Sourced WorkflowOSYour AI Agent Isn't Production Ready. Here's What You're Missing.

Older articles · 32

the dunning kruger ai trap confidence is not competence 2026 06 27ai agents need identity too the production security gap 2026 06 26why ai budgets are getting cut and how to stop it 2026 06 25ai certification for project managers what actually helps 2026 06 24is your ai deployment the next herbalife here s the test 2026 06 24why employers now demand certified ai integrators 2026 06 23ai is eroding your team s skills certification fixes that 2026 06 22why ai deployments fail and what certified integrators do differently 2026 06 21ai raises the engineering bar here s what that means 2026 06 20ai key theft via ide plugins the supply chain gap msps miss 2026 06 19ai agent bankrupted its operator 6 controls that failed 2026 06 18ai agent bankrupted its operator 6 controls that failed 2026 06 17when your oss ai stack disappears overnight 2026 06 16ai agent bankrupted its operator 6 controls that failed 2026 06 15is an ai certification worth it an honest answer 2026 06 15when ai agents go rogue the real cost of skipping guardrails 2026 06 13Free AI Certification That's Actually Free (No Card Required)When AI Hides Its Rules: Claude's Secret GuardrailsAI Certification Compared: Production AI, Cloud Vendor, and GRC Tracksai proof your careerAIDA Certification Study Guide: How to Pass the AI Deployment Associate ExamHow to Certify Your AI Team: A Practical Guide for Engineering and Product LeadersMSP AI Certification Guide: What Your Team Needs and WhyPSF-1: Input Governance — PSF Domain GuidePSF-2: Output Validation — PSF Domain GuidePSF-3: Data Protection — PSF Domain GuidePSF-4: Observability & Monitoring — PSF Domain GuidePSF-5: Deployment Safety — PSF Domain GuidePSF-6: Human Oversight — PSF Domain GuidePSF-7: Security — PSF Domain GuidePSF-8: Vendor Resilience — PSF Domain GuideWhat Is a Certified AI Integrator?
The Production AI Brief

Keep your AI map current

Tool policy changes, new incident records, disclosure signals, and practical next steps. Public evidence, plain English, no hype.

PAI
Production AI Institute

The public record and operating memory for production AI: what changed, what broke, and what the evidence says.

WorkflowOS · open source (MIT)
Navigate
Briefing
OverviewToday's AI briefingPublic record explorer
Check
Check an AI toolRun exposure checkData-use indexIncident registry
The Lab
The LabAI Morality CompassAgent readinessEcosystem assessmentsModel & agent evalsResearch library
Learn
The FrameworkAI Adoption GuideProduction AI Deployment GuideFive-Year Automation RoadmapWorkflowOS · open sourceWorkflow libraryInsights
Act
Ask for disclosureSubmit a correctionSave a watch
Method & trust
How records are madeCorrectionsHow to citeContact
© 2026 Production AI Institute · CC BY 4.0
AboutPrivacyTermsSecurityGovernanceIndependence