Prompt injection defence, input validation, intent classification.
Reference library for production AI
Long-form, versioned reference documents on deploying AI safely. Not blog posts. Citable, maintained, authoritative.
The complete permissions.json schema for Cursor: autoRun.allow_instructions and block_instructions for Auto-review, mcpAllowlist and terminalAllowlist patterns, file location, precedence, and copy-paste examples.
AI tools make everyone sound expert. In production, that overconfidence destroys systems. Here is how to verify AI integrator skills before the damage is done.
Cloudflare's ephemeral agent accounts reveal a governance blindspot most enterprises haven't priced in. Here is what secure agent identity actually requires.
AI budgets are shrinking because deployments lack audit trails, not because AI fails. Learn how production-readiness standards make ROI defensible.
PMs don't need to learn Python. They need AI fluency, governance instincts, and a credential path that starts free at productionai.institute.
CFOs are asking hard ROI questions about AI spend. Here's the production-readiness standard that separates defensible deployments from expensive pilots.
When every resume claims AI expertise, verifiable certification becomes the only trusted hiring signal. Here is what a Certified AI Integrator credential actually proves.
A Nature study confirms AI tool use degrades human expertise. PAI certification verifies the human-in-the-loop competency that prevents production failures.
Enterprise AI budget cuts signal a governance failure, not a technology failure. Learn the three production-readiness gaps driving cost overruns.
Production AI failures aren't skills gaps. They're discipline gaps. Here's what formal engineering rigor looks like - and how to prove your team has it.
The JetBrains plugin attack exposed how MSP AI security reviews miss IDE toolchains. Learn the PSF control gaps and certification steps that close them.
A forensic post-mortem of the DN42 autonomous agent bankruptcy. Six governance controls failed. Here is the exact sequence and how to verify yours.
A forensic post-mortem of the DN42 autonomous agent bankruptcy: six missing governance controls, exact costs, and how to verify them before deployment.
TensorZero archived its repo days after a $7.3M seed round. Here are the 6 supply chain risk signals every certified AI integrator must evaluate.
A forensic post-mortem of the DN42 incident maps six absent governance controls to real costs and shows each failure was preventable.
A balanced cost-benefit analysis for career switchers who are skeptical of credential hype but need a structured, evidence-grounded path into AI.
A forensic breakdown of the DN42 agent incident - the five missing controls, what each failure cost, and how PSF-compliant certification prevents recurrence.
A runaway AI agent bankrupted its operator scanning DN42. Learn the four guardrails every production agent needs and what PSF compliance actually requires.
The bunq penny-transfer exploit exposed how financial AI agents fail without formal security assessment. Here is what certified integrators check before go-live.
PAI's AIDA certification costs nothing and requires no credit card. Learn what it covers, how employers verify it, and when to upgrade to CAAE.
Anthropic's hidden Claude Fable guardrails expose a production AI governance gap. Learn why third-party certification is the only reliable fix.
Weekly AI Data Use Index edition: Google Gemini Apps privacy notice (19 May 2026) for Spark and agentic features, plus OpenAI US privacy policy (18 May 2026). PSF mapping and practitioner actions.
Anthropic confirmed elevated errors across Claude Opus 4.5 through 4.8 and Sonnet 4.6 on June 5, 2026. Verified timeline, PSF vendor-resilience controls, and what production teams should do today.
Federal counsel used Anthropic Claude Console to draft a motion that included fabricated case quotations. Verified facts, PSF mapping, and controls for legal LLM workflows.
An independent PSF assessment of Anthropic Claude Opus 4.8 (May 28, 2026): effort control, dynamic workflows, and mid-task system entries. Strongest on agentic honesty and oversight signals; parallel subagent scale needs deployment guardrails.
Cursor 3.6 (May 29, 2026) routes Shell, MCP, and Fetch calls through a classifier before they run. Our independent assessment: where the gate is real, where allowlists bypass it, and the settings that keep production repos safe.
An independent PSF assessment of Cursor 3.7 browser Design Mode (June 5, 2026): multi-select UI annotation and voice input during agent runs. Stronger human oversight for live UI edits; voice queuing and multi-element scope need explicit production guardrails.
Shared Cursor canvases and embeddable prompt buttons look harmless until they run agents on production repos. PSF assessment of Design Mode, context reports, and the controls to gate before rollout.
An independent PSF assessment of Cursor Enterprise Organizations, Teams, and Groups (GA June 3, 2026): multi-team governance, org-level IdP, and usage analytics. Strong deployment segmentation; permissive-merge rules need explicit policy.
An independent PSF assessment of Google Agent Executor (AX), the open-source distributed agent runtime announced May 20, 2026. Strong on durable execution and human-in-the-loop resumption; governance and data classification remain deployment-layer work.
7,000+ teams will see this in search before they run Cursor SDK 3.7 unattended. Where auto-review helps, where nested subagents inherit dangerous tools, and the permissions.json gates that matter.
An independent PSF assessment of OpenAI Codex CLI 0.134.0 (May 26, 2026): permission profiles, MCP governance, and local history search. Strong sandbox and profile controls; production eval harnesses remain deployment-layer work.
An independent PSF assessment of OpenAI Codex Sites, Annotations, and role-specific plugins (June 2, 2026): 62 enterprise apps and 110 skills for knowledge work. Strong human refinement signals; hosted Sites and SaaS connectors expand data-protection and security scope.
OpenAI's May 29, 2026 multi-service outage broke ChatGPT, API, and login for hours. Verified timeline from the status page, PSF vendor-resilience controls, and what to test before the next failure.
An independent PSF assessment of OpenAI GPT-5.5, GPT-5.4, and Codex on Amazon Bedrock (GA June 1, 2026): AWS-native governance, regional inference, and Codex routing through Bedrock. Strong data protection and audit controls; human approval for autonomous Codex remains deployment-layer work.
Empirical PAI Lab scan of 20 public AI agent repositories against PSF evidence signals. Methodology PAI-ARI-2026.1, domain coverage, limits, and practitioner actions.
Cursor 3.5 Automations run cloud agents on a schedule — including no-repo ops agents. Where scheduling UX is strong, where data scope and merge policy still need explicit gates.
Google launched Cross-Corpus Retrieval powered by Agentic RAG in Gemini Enterprise Agent Platform on June 5, 2026. Production impact, PSF control implications, and what to do today.
NVIDIA released Nemotron 3.5 Content Safety on June 4, 2026 with multimodal checks, custom enterprise policies, and auditable THINK traces. Production impact and PSF control implications.
OpenAI added inline moderation scores to the Responses API and Chat Completions API on June 4, 2026. Production impact, PSF output-validation controls, and what to do today.
Agentic design patterns are the building blocks of every production AI system. This guide translates all 21 patterns — from prompt chaining to multi-agent orchestration — into plain language for domain experts who are not software engineers.
A 32-item production readiness checklist for AI agents mapped to the eight PSF domains. Use it to decide whether your agent is ready to ship, needs hardening, or should stay in staging.
An independent comparison of AI certification paths for 2026: Production AI Institute PSF credentials, ISACA AAISM/AAIA/AAIR, AWS, Microsoft Azure, and Google Cloud. Matrix, decision guide, and what production deployments require.
A practitioner's comparison of the three leading AI governance frameworks - Production Safety Framework (PSF), ISO 42001, and NIST AI RMF. What each covers, where they overlap, and how to choose.
Complete study guide for the PAI AIDA certification. What the exam covers, how it is structured, the topics that appear most frequently, and how to prepare in under 2 hours.
An independent assessment of Amazon Bedrock Agents against the eight domains of the Production Safety Framework — enterprise-grade strengths, critical gaps, and what practitioners must configure for safe AI deployment.
Anthropic suspended Claude Fable 5 on June 13, 2026, hours after Production AI Institute completed a full CompassEval sweep of the model. Why independent evaluation matters, and what the suspension reveals about frontier AI model governance.
An independent assessment of Microsoft AutoGen / AG2 against the eight PSF domains. AutoGen was designed for research, not production — this assessment documents what practitioners must add before enterprise deployment.
A PSF-based comparison of the four major agent frameworks for production AI deployment. Independent assessment of safety profiles, operational maturity, and ecosystem requirements for each.
An independent assessment of Anthropic Claude Sonnet 4.6 against the eight domains of the Production Safety Framework. Strongest current model on human oversight and refusal calibration; deployment-layer data protection still required.
An independent assessment of Composio against the eight domains of the Production Safety Framework. Where it satisfies requirements, where it leaves gaps, and what practitioners must add.
An independent assessment of CrewAI against the eight PSF domains. Multi-agent architectures amplify every safety gap — this assessment explains where CrewAI leaves practitioners exposed and what must be added before production deployment.
The Cursor SDK (released April 2026) gives developers programmatic access to the same agent runtime that powers the Cursor IDE. This assessment evaluates its production safety profile against the PSF — with particular attention to the data protection and security requirements of agents with filesystem and MCP access.
Cursor SDK lets you embed the same agent runtime that powers Cursor into CI/CD pipelines, event-driven workflows, and your own products. This guide shows three production deployment patterns with full PSF Domain compliance controls applied.
Dify v1.14.2 tightens tenant isolation, tracing, and deployment defaults. A PSF assessment for teams running production workflows.
Independent PSF assessment of DSPy (Stanford). The optimisation-first framework with the strongest structured output guarantees — and the widest gap between research elegance and production safety.
Compliance and governance guide for AI systems in energy, utilities, and critical infrastructure. Covers NERC CIP, IEC 62443, NIS2, NIST CSF, OT/IT convergence, and PSF domain mapping for safety-critical deployments.
PSF assessment of no-code agent builders Flowise and LangFlow. Fast to deploy, fastest to expose gaps. What every enterprise evaluation needs to know before going to production.
An independent assessment of Google Gemini 1.5 Pro against the eight domains of the Production Safety Framework. Strong on long-context handling and Google Cloud integration; security posture and prompt-injection resistance remain mid-cohort.
An independent assessment of OpenAI GPT-4.1 against the eight domains of the Production Safety Framework. Where the model is production-strong, where deployment-layer controls remain mandatory, and how the score is derived.
Which guardrails tool closes which PSF D1 and D2 gaps? Independent comparison of the three major production AI safety libraries.
Independent PSF assessment of Haystack by deepset. The RAG-native framework with the strongest production deployment story of any Python agent framework.
A practitioner guide to deploying AI in healthcare settings. Covers HIPAA compliance, FDA AI/ML guidance, clinical decision support safety, and PSF domain mapping for regulated clinical environments.
A step-by-step guide to building a certified AI team — which credentials matter, how to sequence them, and how to build a team-wide certification programme without disrupting delivery.
A structured guide to defining, documenting, and enforcing the operational boundaries of production AI systems through formal behaviour contracts.
A practitioner guide to deploying AI in HR and employment contexts. Covers EEOC guidance on AI hiring tools, NYC Local Law 144 bias audits, EU AI Act high-risk classification, employee monitoring AI, and PSF domain mapping.
A practical framework for deciding where human oversight belongs in AI workflows, and how to design it so it actually works.
An independent assessment of LangChain and LangGraph against the eight domains of the Production Safety Framework — where the ecosystem is strong, where it leaves gaps, and what companion tooling practitioners should add.
LangGraph 1.2.1 tightens stream handling and message hygiene in a framework built for stateful agents. A PSF assessment for production teams.
Independent PSF Domain 4 comparison of the three major production AI observability platforms. Which tool gives you the trace visibility your production deployment actually needs?
A practitioner guide to deploying AI in legal and government settings. Covers the EU AI Act high-risk requirements, FedRAMP/FISMA compliance, algorithmic bias in justice systems, and full PSF domain mapping.
An independent assessment of Meta Llama 3.1 70B in self-hosted production deployments against the Production Safety Framework. Strong on data sovereignty; significant deployment-layer investment required for security, observability, and oversight.
An independent assessment of LlamaIndex against the eight domains of the Production Safety Framework — how the data framework performs for enterprise RAG and agent deployments, critical gaps, and what practitioners must add.
Independent Production Safety Framework assessment of Microsoft Semantic Kernel. Domain-by-domain analysis for enterprise .NET and Python AI deployments.
MSPs deploying AI for clients need staff who can assess, deploy, and govern AI systems safely. This guide covers the certifications that matter, the skills gaps most MSPs face, and how to build a certified AI practice.
An independent assessment of n8n against the eight domains of the Production Safety Framework — how the workflow automation platform performs for enterprise AI deployments, where gaps exist, and what practitioners must add.
An independent assessment of the OpenAI Agents SDK against the eight domains of the Production Safety Framework — how the SDK performs for production agent deployments, where gaps exist, and what practitioners must add.
Practical guide to improving AI discovery across ChatGPT, Claude, Perplexity, and Gemini with llms.txt, crawl controls, and structured content design.
A PSF D3/D4 assessment of the three major vector databases. Data residency, PII in embeddings, audit logging, managed vs self-hosted tradeoffs — everything that matters for production RAG deployments.
How financial services firms map PSF requirements to FCA, SEC, MiFID II, DORA, and SR 11-7 obligations. Domain-by-domain guidance for regulated AI deployments.
PSF compliance means your AI system meets the eight domains of the Production Safety Framework. This guide explains what each domain requires, how compliance is assessed, and which certifications demonstrate it.
The complete practitioner guide to PSF Domain 1 — prompt injection defence, input classification, sanitisation, schema enforcement, and the companion tooling that closes the gap every framework leaves open.
The complete practitioner guide to PSF-2 Output Validation — output contracts, schema enforcement, confidence thresholds, semantic validation, and how to design output pipelines that catch model failures before they reach users.
D3 is a Gap for every major agent framework. This guide explains why, maps the full data protection threat surface for production AI, and documents what practitioners must implement themselves.
The complete practitioner guide to PSF-4 Observability — what to log, how long to retain it, alerting patterns, drift detection, and the specific requirements that distinguish meaningful AI observability from basic request logging.
The complete D5 implementation guide. Covers model version control, canary deployment patterns, rollback triggers, predetermined change control plans, and the deployment anti-patterns that cause silent production failures.
The complete guide to PSF Domain 6 — when human oversight is required, how to design effective escalation, autonomy level frameworks, and how to avoid the compliance theatre trap.
The complete D7 implementation guide. Covers the AI-specific threat model, prompt injection at infrastructure scale, model supply chain attacks, adversarial inputs, and the security controls that close gaps every framework leaves open.
The complete D8 implementation guide. Covers vendor lock-in risk taxonomy, model deprecation planning, multi-vendor architecture patterns, SLA benchmarking, exit strategy documentation, and the resilience controls every framework leaves out.
Everything your AI system accepts is a potential attack surface. PSF-1 covers prompt injection defence, input validation, intent classification, and the sanitisation practices that prevent malicious inputs from reaching your models.
AI outputs must be validated before they act on the world. PSF-2 covers output schema enforcement, content filtering, hallucination detection, structured output contracts, and graceful degradation when outputs fail validation.
AI systems process sensitive data at unprecedented scale. PSF-3 covers PII handling, consent chains, data residency, retention obligations, and the specific data governance requirements that apply when personal data flows through AI inference pipelines.
You cannot manage what you cannot measure. PSF-4 covers inference logging, drift detection, quality scoring pipelines, cost and latency monitoring, and the alerting architecture that keeps production AI systems visible and governable.
Safe deployment of AI systems requires more than passing a pre-production test. PSF-5 covers canary releases, rollback procedures, circuit breakers, environment parity, and the operational controls that keep new AI deployments from becoming production incidents.
Human oversight is not a checkbox — it is a designed system capability. PSF-6 covers autonomy level assignment, escalation path design, human review queue architecture, override mechanisms, and the organisational practices that keep human judgment genuinely in the loop.
AI systems introduce novel attack surfaces that conventional security controls were not designed for. PSF-7 covers AI-specific threat modelling, API key management, model access controls, adversarial robustness, and the security practices that protect production AI infrastructure.
Every production AI system that depends on a third-party model provider is one API change, deprecation notice, or outage away from failure. PSF-8 covers abstraction layer design, vendor portability, multi-provider architecture, deprecation response planning, and the governance practices that prevent vendor lock-in from becoming a production crisis.
Pre-validated, PSF-compliant stack combinations for every major agent framework. If you're using LangGraph, here's the full stack. If you're using CrewAI, here's what you need to add.
Independent PSF assessment of Pydantic AI. Type-safe structured outputs as a first principle — with meaningful gaps in human oversight and deployment safety.
A practitioner guide to deploying AI in retail and e-commerce. Covers personalisation dark patterns (EU DSA, FTC), recommendation system bias, fraud detection AI, customer service chatbots, and full PSF domain mapping for retail deployments.
Temporal Python SDK 1.27.2 adds explicit DNS-based load balancing behavior for hostname targets. A PSF assessment for production teams.
The agent operator is the most important new role in enterprise technology. They don't need to be engineers. They need MCPs, CLIs, file writing, agents.md fluency, and business acumen. No CS curriculum teaches this yet.
A new class of AI deployment is emerging: agents that live inside existing tools rather than as standalone interfaces. Gmail agents, IDE agents, browser agents. The PSF requirements for ambient agents differ from isolated API deployments in ways that most safety frameworks have not addressed.
If morality can be grounded in physics rather than religion, and if a machine could one day be the most powerful moral agent on earth, it needs a compass. The philosophical case for The Compass framework.
Multi-agent architectures don't just add PSF gaps — they multiply them. This reference documents blast radius amplification, agent-to-agent trust, shared context risks, and safe multi-agent architecture patterns.
A taxonomy of how AI systems fail in production environments, with diagnostic patterns and mitigation strategies for each mode.
The CPU dominated computing for 40 years. The GPU displaced it for AI. Now Karpathy says it is happening again — and whoever owns the model wins the third era. Here is what that means for your career and your organisation.
An independent assessment of Trinity by Ability.AI against the eight PSF domains. Trinity is the most governance-forward agent runtime we have assessed — self-hosted, sovereign, and security-audited. This article explains what it covers natively and where MSPs and practitioners still need to add controls.
The Certified AI Integrator credential recognises organisations that can deploy, govern, and maintain production AI systems to a verified standard. This guide explains what it covers, who it is for, and how assessment works.
Clear definition of a production AI system, what is excluded, and the PSF readiness criteria required before real-world deployment.
Production AI means AI running in live operations with governance, observability, and accountable outcomes. A practitioner definition, maturity model, and links to the PSF standard.
Andrej Karpathy built an entire GPT in 200 lines of pure Python. No libraries. No frameworks. This is what those 200 lines actually tell you about how language models work — explained for professionals who are not software engineers.
A practical guide to EU AI Act compliance for teams deploying AI systems in production. Covers risk classification, conformity assessment, and ongoing obligations.
WorkflowOS is the MIT-licensed reference implementation of the Production Safety Framework - a free PSF workflow designer you can use hosted or self-host for client work.
You built your first AI agent. It works in testing. Before you ship it to real users or real business processes, read this: the eight things every production AI agent needs that tutorials never teach.
Domain reference guides
Deep reference for all eight Production Safety Framework domains. Aligned with examination content. Licensed CC BY 4.0.
Schema enforcement, hallucination detection, confidence gating.
PII handling, consent chains, data residency, vector store risk.
Inference logging, quality scoring, drift detection and alerting.
Canary releases, rollback procedures, circuit breakers.
Autonomy levels, escalation design, blind review sampling.
Threat modelling, adversarial robustness, red-teaming architecture.
Abstraction layers, version pinning, fallback provider design.
Library index
The full PAI research library, grouped by type.