CyberMindSpace LABSTalk to us ↗
CYBERMINDSPACE LABS / RESEARCH

Security, as an experimental science.

A research program for programmable, reproducible and measurable security environments.

PROPOSED / RESEARCH
RESEARCH 001 / THESIS 0.2

Programmable.
Reproducible.
Measurable.

Read the complete thesis ↗

Can security intent survive the journey from a scenario definition to a deployed environment and a defensible verdict?

That is the central question. Our proposed program focuses first on a bounded AI-agent environment, where tool permissions, protected resources and evidence can be inspected together.

No experimental results are claimed. The thesis distinguishes established work, proposed contributions and experiments that could falsify the hypotheses.

01

AI-system security

Models, agents, retrieval and tool permissions.

02

Dynamic environments

Typed scenarios and validated initial state.

03

Scenario generation

Constrained transformations that preserve intent.

04

State-based evaluation

Security outcomes supported by independent evidence.

05

Security telemetry

Correlated events with explicit provenance.

OPEN QUESTIONS

Questions worth testing.

RQ / 01Can security intent survive environment generation?+
Hypothesis

A typed scenario contract can preserve a defined security property across constrained changes.

Technical challenge

A deployable configuration can still remove the attack path or invalidate the grader.

Planned experiment

Generate controlled variants with independent benign and vulnerability witnesses.

Measurement

Valid-run fraction, missed semantic defects and total authoring/repair time.

RQ / 02Can we measure AI security across system boundaries?+
Hypothesis

Independent state evidence can distinguish model statements from actual tool effects.

Technical challenge

A model can refuse in its answer after already performing an unauthorized action.

Planned experiment

Compare output grading, existing state-aware evaluation and the proposed contract evaluator.

Measurement

Precision, recall, inconclusive rate and legitimate task completion.

RQ / 03What does reproducible mean for a live model?+
Hypothesis

Separate fixture replay from equivalent state reconstruction and statistical repetition.

Technical challenge

Provider changes and asynchronous scheduling prevent universal determinism.

Planned experiment

Repeat pinned and live-model conditions while changing one dependency at a time.

Measurement

State equivalence, verdict agreement, outcome distributions and reproduction time.

RQ / 04Can incomplete telemetry produce honest verdicts?+
Hypothesis

Evidence obligations can prevent missing observations from becoming false passes.

Technical challenge

Events can be lost, delayed, duplicated or forged by a compromised target.

Planned experiment

Inject evidence faults and ablate sensors against independent resource witnesses.

Measurement

False-pass rate, erroneous graph edges, evidence loss and detection delay.

R&D ROADMAP / NO COMMITTED DATES

Build the foundation.
Earn the next layer.

PHASE / 01

Foundation

Basic cybersecurity and prompt-injection labs. Classroom workflows. Environment lifecycle.

IN DEVELOPMENT
PHASE / 02

Programmable environments

Scenario-as-code, dynamic provisioning, isolation and telemetry.

RESEARCH
PHASE / 03

AI-security infrastructure

Agents, RAG, MCP, tools and validated automated evaluation.

RESEARCH
PHASE / 04

Environment intelligence

Constrained scenario generation, event graphs and reproducible experiments.

LONG-TERM RESEARCH
BUILD THE RESEARCH WITH US

The next experiment
starts with a question.

Register your interest ↗