CyberMindSpace LABSTalk to us ↗
CYBERMINDSPACE LABS / RESEARCH / TELEMETRY

See the attack as a system.

Research on cross-boundary events, provenance and uncertainty.

PROPOSED / RESEARCH
EVIDENCE INSPECTOR / RUN FIXTURE-042SYNTHETIC RECEIPTS
EVENT / e01 · 1 OF 5

Initial state verified

Tenant A · support_reader · synthetic record CRM-042; public summary task available.

event_id: e01
parent_event_id: none
source: harness
run_id: fixture-042

Parent links are declared by this fixture. In production, each link must be supported by trusted propagation and resource identifiers; timestamps alone are insufficient.

Telemetry and security event graphs

Adopt OpenTelemetry-compatible tracing and map relevant infrastructure events to existing security schemas where practical. The GenAI conventions are under development, so record schema versions and maintain explicit adapters. OCSF provides a vendor-neutral security schema; neither framework establishes causality by itself.

A proposed event envelope contains experiment_id, environment_id, event_id, source_id, source_sequence, observed_time, ingested_time, trace_id, parent_id, principal, action, resource, policy_decision, result, schema_version and evidence_reference. Content-bearing payloads are minimized and access-controlled. Synthetic canaries support evaluation without collecting real customer secrets.

Graph edges must distinguish observed invocation, explicit data provenance and inferred temporal association. Two events occurring close together do not prove one caused the other. Clock skew, retries and asynchronous work require source sequences, correlation identifiers and uncertainty labels. A graph may contain gaps; displaying them is part of measurement integrity.

Do not claim access to hidden model reasoning. Record observable prompts, outputs, declared actions and tool behavior. An agent's self-reported rationale is neither a faithful internal trace nor authoritative evidence of why an action occurred.

The research question is whether this evidence structure improves cross-boundary diagnosis and evaluator reliability compared with ordinary correlated traces. A compelling animation is not a scientific contribution.

Reproducibility

Separate three guarantees. Fixture replay reuses recorded model and service responses to debug the harness. Environment reproduction reconstructs equivalent initial resources and policies under a declared equivalence relation. Live repetition reruns the experiment against a model or service and measures an outcome distribution.

Do not promise bit-identical live model output. A seed, temperature setting or version label may not fix provider implementation, scheduling or backend changes. Record those limits. A replayed successful attack demonstrates that recorded events can be replayed, not that the same attack succeeds today.

Snapshots also need fresh per-session identities, secrets and entropy where uniqueness is required. Reproducing experimental conditions must not clone credentials across learners. Canonical manifests should distinguish stable experimental fields from deliberately regenerated operational fields.

Version every evaluator and retain the old rubric with its evidence. A later scorer can reanalyze an earlier bundle, but the resulting finding must be labeled a reanalysis. Compare failure distributions and security-state transitions, not only exact transcript hashes.

Open research questions

What is the smallest scenario language that captures meaningful security properties without becoming a general-purpose programming language? Which properties can be validated before execution, and which require independent runtime witnesses? How much evidence is enough to distinguish a prevented action from an unobserved action?

Can scenario mutation preserve attack feasibility while changing superficial details? How should statistical reproduction be reported when a provider cannot pin its backend? Can evaluator portability survive differences between synthetic and real authorization systems? What is the smallest counterexample that explains a failed security property to both a learner and an engineer?

When does additional infrastructure realism improve measurement enough to justify its cost? Can a constrained domain outperform integration of existing tools, and does that improvement persist when another team authors the scenarios?

Full thesis and source references ↗
BUILD THE RESEARCH WITH US

The next experiment
starts with a question.

Register your interest ↗