H11I official logo

H11I intelligence agent

H11-EVALUATOR

Benchmarking, capability measurement, behavioral evaluation, regression, and system-assurance intelligence

Unique intelligence

The evaluation intelligence that turns capability claims into representative, adversarial, longitudinal, contamination-aware, and reproducible promotion evidence.

Intelligence planeMeta Intelligence
Agent categoryevaluation-intelligence
Sovereign modelH11-SIM

Role in the council

Benchmarking, capability measurement, behavioral evaluation, regression, and system-assurance intelligence

This specialist contributes to adaptive H11I councils while evidence, authority, verification, and execution remain separated by platform governance.

Platform engine contract

H11-SIM activates H11-EVALUATOR when the present user intent requires benchmarking, capability measurement, behavioral evaluation, regression, and system-assurance intelligence; the result is used as one verified contribution inside the 110-agent council rather than as a standalone chatbot reply.

Runtime pipeline

  1. lock the relevant user intent and success condition
  2. build the agent-specific task frame
  3. requirement-to-metric mapping
  4. evaluation dataset design
  5. trajectory and outcome scoring
  6. regression and contamination analysis
  7. emit typed artifacts with uncertainty and failure flags
  8. hand off to H11I-VERITAS / H11I-ZENITH for release synthesis

Typed artifacts

Quality gates

Authority boundary: bounded specialist analysis and typed recommendation; no independent external authority

Failure recovery: Return the failed gate, preserve the strongest verified partial result, and request escalation/revision instead of inventing certainty.

Source backing