Knowledge Portal · engineering documentation

Skip to content

Benchmark

Status: Canonical
Owner: PL-006
Last updated: 2026-07-05


Dataset-driven evaluation of agent behavior — read-only with respect to production runtime.


Principles

RuleDetail
IsolationBenchmark code paths do not mutate live orchestrator config
ReproducibilityDataset version pinned in report
RegressionEvery release compares against prior baseline

Flow

mermaid
flowchart LR
    DS[Ground truth dataset] --> RUN[Benchmark runner]
    RUN --> RPT[Report JSON]
    RPT --> ACC[Acceptance Lab]

Breadcrumbs: Home → AI → Benchmark

ZAIXOS Knowledge Portal — public engineering docs at /docs · Staff operations at /admin