State-based evaluation
Make state-based evaluation a core capability, while explicitly recognizing existing research such as AgentDojo. The proposed differentiator is validated correspondence between the scenario property, observation points and evaluator across controlled environment changes.
For disclosure, compare a synthetic protected value or a defined transformation of it with evidence at an unauthorized sink. For a forbidden write, check the destination service's durable state and authenticated audit record. For a privilege change, compare effective authorization before and after the run. For remediation, verify restored state and rerun the attack plus the legitimate task.
Exact matching misses paraphrase and encoding; semantic judges can misclassify or be influenced by adversarial text. Use deterministic checks where ground truth is available. Treat model-based assessments as fallible supporting signals with separately measured agreement, and never give an evaluation model authority to change the environment or its own rubric.
Measure security and legitimate task completion separately. A system that refuses every request may block an attack while becoming useless. Report attempted action, permitted action and completed effect as distinct events, and document whether a test evaluates confidentiality, integrity, availability or functional utility.
Experimental methodology and decision gates
Pre-register hypotheses, primary endpoints, resource budgets, exclusions and stopping rules before confirmatory runs. Keep development scenarios separate from held-out evaluation scenarios; split by scenario family when testing generalization. Preserve infrastructure failures in operational denominators and report conditional security results separately.
Use paired comparisons where the same scenario can run under both methods. Treat scenarios, rather than repeated prompts alone, as units when estimating generalization. Report uncertainty intervals and effects, use cluster-aware resampling where appropriate, and account for multiple comparisons. Pilot counts above are planning values, not a statistical power justification.
Proposed engineering gates are at least 95% readiness success across the supported pilot corpus, no false pass on deliberately missing required evidence, and verified cleanup for every pilot run. A proposed research gate is at least a 25% reduction in median authoring-and-repair time against the integrated baseline without worse classification quality. These thresholds are managerial choices to preregister and revise before validation, not literature-derived standards.
A zero-escape pilot result must be reported with the tested actions and exposure, never as proof of perfect isolation. If cost per useful experiment or integration effort exceeds a viable commercial budget, narrow the supported environment family even when technical metrics look good.
Limitations
This is a proposed program, not a proven platform. The source review is targeted and cannot establish an exhaustive market or patent landscape. Product documentation is evidence of advertised capabilities, not independently verified effectiveness. Academic comparisons refer to the cited versions and their particular threat models.
A constrained support-agent family cannot establish generality across operating-system compromise, Active Directory, arbitrary cloud services or autonomous multi-agent systems. Simulated APIs can hide real authorization and consistency behavior. Production model providers can change. Telemetry can be compromised, incomplete or too costly to retain in full.
A platform that constructs its own scenarios and judges its own evidence can share a common error across both components. Independent resource witnesses, external review and reproductions are essential. A well-formed experiment can still measure the wrong security property.