CyberMindSpace LABSTalk to us ↗
CYBERMINDSPACE LABS / RESEARCH / EVALUATION

A verdict must have evidence.

Research on state-based security evaluation and honest inconclusive outcomes.

PROPOSED / RESEARCH
RQ / 01Can security intent survive environment generation?+
Hypothesis

A typed scenario contract can preserve a defined security property across constrained changes.

Technical challenge

A deployable configuration can still remove the attack path or invalidate the grader.

Planned experiment

Generate controlled variants with independent benign and vulnerability witnesses.

Measurement

Valid-run fraction, missed semantic defects and total authoring/repair time.

RQ / 02Can we measure AI security across system boundaries?+
Hypothesis

Independent state evidence can distinguish model statements from actual tool effects.

Technical challenge

A model can refuse in its answer after already performing an unauthorized action.

Planned experiment

Compare output grading, existing state-aware evaluation and the proposed contract evaluator.

Measurement

Precision, recall, inconclusive rate and legitimate task completion.

RQ / 03What does reproducible mean for a live model?+
Hypothesis

Separate fixture replay from equivalent state reconstruction and statistical repetition.

Technical challenge

Provider changes and asynchronous scheduling prevent universal determinism.

Planned experiment

Repeat pinned and live-model conditions while changing one dependency at a time.

Measurement

State equivalence, verdict agreement, outcome distributions and reproduction time.

RQ / 04Can incomplete telemetry produce honest verdicts?+
Hypothesis

Evidence obligations can prevent missing observations from becoming false passes.

Technical challenge

Events can be lost, delayed, duplicated or forged by a compromised target.

Planned experiment

Inject evidence faults and ablate sensors against independent resource witnesses.

Measurement

False-pass rate, erroneous graph edges, evidence loss and detection delay.

State-based evaluation

Make state-based evaluation a core capability, while explicitly recognizing existing research such as AgentDojo. The proposed differentiator is validated correspondence between the scenario property, observation points and evaluator across controlled environment changes.

For disclosure, compare a synthetic protected value or a defined transformation of it with evidence at an unauthorized sink. For a forbidden write, check the destination service's durable state and authenticated audit record. For a privilege change, compare effective authorization before and after the run. For remediation, verify restored state and rerun the attack plus the legitimate task.

Exact matching misses paraphrase and encoding; semantic judges can misclassify or be influenced by adversarial text. Use deterministic checks where ground truth is available. Treat model-based assessments as fallible supporting signals with separately measured agreement, and never give an evaluation model authority to change the environment or its own rubric.

Measure security and legitimate task completion separately. A system that refuses every request may block an attack while becoming useless. Report attempted action, permitted action and completed effect as distinct events, and document whether a test evaluates confidentiality, integrity, availability or functional utility.

Experimental methodology and decision gates

Pre-register hypotheses, primary endpoints, resource budgets, exclusions and stopping rules before confirmatory runs. Keep development scenarios separate from held-out evaluation scenarios; split by scenario family when testing generalization. Preserve infrastructure failures in operational denominators and report conditional security results separately.

Use paired comparisons where the same scenario can run under both methods. Treat scenarios, rather than repeated prompts alone, as units when estimating generalization. Report uncertainty intervals and effects, use cluster-aware resampling where appropriate, and account for multiple comparisons. Pilot counts above are planning values, not a statistical power justification.

Proposed engineering gates are at least 95% readiness success across the supported pilot corpus, no false pass on deliberately missing required evidence, and verified cleanup for every pilot run. A proposed research gate is at least a 25% reduction in median authoring-and-repair time against the integrated baseline without worse classification quality. These thresholds are managerial choices to preregister and revise before validation, not literature-derived standards.

A zero-escape pilot result must be reported with the tested actions and exposure, never as proof of perfect isolation. If cost per useful experiment or integration effort exceeds a viable commercial budget, narrow the supported environment family even when technical metrics look good.

Limitations

This is a proposed program, not a proven platform. The source review is targeted and cannot establish an exhaustive market or patent landscape. Product documentation is evidence of advertised capabilities, not independently verified effectiveness. Academic comparisons refer to the cited versions and their particular threat models.

A constrained support-agent family cannot establish generality across operating-system compromise, Active Directory, arbitrary cloud services or autonomous multi-agent systems. Simulated APIs can hide real authorization and consistency behavior. Production model providers can change. Telemetry can be compromised, incomplete or too costly to retain in full.

A platform that constructs its own scenarios and judges its own evidence can share a common error across both components. Independent resource witnesses, external review and reproductions are essential. A well-formed experiment can still measure the wrong security property.

Full thesis and source references ↗
BUILD THE RESEARCH WITH US

The next experiment
starts with a question.

Register your interest ↗