Judgment Engineering

The human looping through the system.

Use Judgment Engineering to develop and apply human judgment throughout AI-enabled work—so people remain able to frame the problem, challenge assumptions, evaluate evidence, direct action, and learn from the consequences.

One discipline. Four functions.
  1. FORGEDevelop human judgment
  2. LOOPEngage throughout the system
  3. STEERGovern AI-enabled work
  4. JUDGEEvaluate what AI produces

↺ Feed learning back into FORGE

Judgment Engineering is the deliberate discipline of designing, developing, and applying human judgment throughout AI-augmented work. It brings subject-matter expertise, experience, critical thinking, ethics, and accountability into the decisions people make with AI.

Human judgment belongs wherever the work calls for it. Across conversations with large language models, individual agents, and agentic systems, people need recurring opportunities to question, redirect, authorize, verify, escalate, and stop. Each outcome should also help them learn.

Prompt design and the INSPIRE methodology help people communicate their intentions to AI. Judgment Engineering addresses the broader decision environment: the problem being framed, the evidence being trusted, the authority to act, and responsibility for what happens next.

The governing principle

AI can contribute intelligence, analysis, generation, recommendations, and action.

Humans retain responsibility for purpose, wisdom, authorization, ethics, and consequences.

The aim is appropriate reliance: knowing when to accept, modify, reject, investigate, seek subject-matter expertise, escalate, or stop an AI-supported course of action.

FORGEHow human judgment is developed

Develop the capability to make sound decisions through practice, challenge, feedback, reflection, and experience. This includes productive friction: taking time to test an easy answer, consider disagreement, and learn from what happened.

F — Frame the Problem

Develop the discipline to define the real problem, decision, stakeholders, criteria, constraints, stakes, and decision rights before seeking an answer.

O — Open Alternatives

Resist premature certainty by deliberately exploring different explanations, perspectives, scenarios, approaches, and possible courses of action.

R — Review Evidence

Build the habit of distinguishing claims from evidence, testing sources and reasoning, recognizing uncertainty, and determining what deserves belief.

G — Govern the Decision

Practice making and defending decisions under uncertainty while balancing expertise, values, risk, consequences, accountability, and AI involvement.

E — Evolve Through Feedback

Strengthen judgment through experience, mistakes, reflection, disagreement, feedback, recalibration, and repeated exposure to consequential decisions.

FORGE is the development mechanism. Judgment grows through subject-matter expertise, deliberate practice, reflection, and experience.

LOOPWhere human judgment enters the system

Identify the recurring points where human judgment needs to enter and shape the work. Place checkpoints, feedback, and intervention where uncertainty, authority, or consequences make them necessary.

L — Locate the Judgment Points

Identify where uncertainty, consequences, expertise, ethics, competing priorities, or decision authority require meaningful human judgment.

O — Orient the System

Establish purpose, intent, context, stakeholders, criteria, constraints, values, and desired outcomes before the system determines direction.

O — Oversee and Intervene

Remain able to question, challenge, correct, redirect, constrain, approve, escalate, or stop the system as conditions and evidence change.

P — Process the Feedback

Examine outcomes and consequences, capture what was learned, recalibrate assumptions and standards, and feed that learning into the next iteration.

LOOP is the system architecture. Human judgment repeatedly enters wherever it adds value or protects against consequential error.

STEERHow humans govern AI-enabled work

Maintain direction, agency, and accountability while the work unfolds. These behaviors help people stay engaged as evidence, circumstances, and recommendations change.

S — Set Intent

Define what the system is trying to accomplish, why the outcome matters, what success means, and what must remain under human authority.

T — Test Assumptions

Surface and challenge assumptions made by both humans and AI before they quietly become the foundation for recommendations or actions.

E — Evaluate Evidence

Examine the quality, relevance, completeness, provenance, uncertainty, and competing interpretations of the evidence supporting a conclusion.

E — Exercise Judgment

Integrate evidence with expertise, experience, ethics, context, stakeholder impact, and consequences to determine what should actually happen.

R — Recalibrate

Adjust the decision, boundaries, instructions, criteria, confidence, or level of AI reliance when evidence, conditions, or understanding change.

STEER is the governance behavior. AI provides capability and momentum; humans remain responsible for direction.

JUDGEHow AI-supported recommendations are evaluated

Examine an AI-supported output before relying on it. Use these questions to separate a convincing presentation from evidence and reasoning that deserve to influence action.

J — Jurisdiction

Determine what decision is actually being made, whether AI belongs in it, who has authority to act, and who remains accountable for the result.

U — Unpack the Output

Separate claims, assumptions, evidence, interpretations, recommendations, uncertainties, and omissions instead of treating polished output as inherently sound.

D — Demand Evidence and Dissent

Require supporting evidence, alternatives, counterarguments, missing perspectives, failure modes, and information that could disconfirm the recommendation.

G — Grade the Reasoning

Evaluate whether the reasoning is logical, sufficiently supported, contextually appropriate, ethically defensible, and proportionate to the risk.

E — Exercise Human Judgment

Choose to accept, modify, reject, investigate further, seek expertise, escalate, or stop the AI-supported course of action—and own the decision.

JUDGE is the evaluation discipline. AI produces possibilities and recommendations; humans determine what deserves trust, reliance, action, and publication.

A continuous learning cycle

These four functions work together. FORGE develops the human capability. LOOP places it throughout the system. STEER guides the work as it happens. JUDGE evaluates what the system produces. The learning from those decisions strengthens judgment for the next cycle.

FORGE → LOOP → STEER → JUDGE → FORGE

People retain the ability to frame, govern, override, escalate, decide, learn, and own consequences throughout the system. Human agency becomes part of how the work is designed.

The human capabilities behind the system

Judgment Engineering depends on capabilities people develop and carry from one setting to another:

  • Critical reasoning
  • Metacognition
  • Subject-matter expertise
  • Experience
  • Leadership judgment
  • Ethical reasoning
  • Communication
  • Systems thinking
  • AI awareness, literacy, and fluency
  • Accountability
  • Human agency

These capabilities help people govern what AI tools produce and do, while continuing to develop their own judgment.

Bring Judgment Engineering into your work.

Where should human judgment enter your AI-enabled work? Let’s explore the decisions, capabilities, and checkpoints that matter for you and your team.

Start a conversation

johnggartin@gmail.com

John Gartin
Better decisions. Stronger humans. Better systems.