Off-distribution
A substation control loop or a correspondent-banking settlement is nothing like the text a model was pretrained on.
Cyberange · Data · For frontier models
Frontier models learn cyber-physical operations from web text and generic logs. It's a world they've never touched. We run real attack-and-defend campaigns on live OT and critical-infrastructure ranges, then turn the signals into training data and benchmarks a model can learn from.
Power systems · Interbank payments & SWIFT · Rail signalling — more on the way
The gap
Models are strong on code and prose because the internet is saturated with both. Cyber-physical operations are the opposite. The telemetry lives on air-gapped networks, incident data is confidential, and nobody publishes what a real ICS attack looks like from initial access to physical impact. So a model asked to reason about a tripping relay, a fraudulent MT103, or a mis-set signalling interlock is guessing by analogy — from text that describes these systems but never records one operating.
We close that gap at the source. Rather than scrape for events that were never recorded, we stage them on real plant, with expert operators on one side and a live adversary on the other.
A substation control loop or a correspondent-banking settlement is nothing like the text a model was pretrained on.
The rare real logs that surface have no ground-truth: which packet was the attack, which action was the fix.
No public source connects network, control, operator and physical layers through a single incident.
The engine
One pipeline turns an attack-and-defend campaign into data a model can train and be measured on. Every step is instrumented, so nothing is reconstructed after the fact.
Red and blue teams operate real scenarios on our physical and virtual ranges: a substation, an interbank payment network, a metro line.
Every layer is captured at once: packet, protocol, process variable, operator action, physical sensor. All on the same clock.
Domain experts annotate against attack ground-truth: technique, phase, intent, the correct response, the actual outcome.
Signals become model-ready: instruction pairs, agent trajectories, preference data, graded eval tasks, all in your schema.
Versioned training corpora plus a sealed, held-out benchmark. Every record ships with its provenance and licence.
Signals in, formats out
What we capture
Full pcap plus decoded Modbus/TCP, DNP3, IEC 60870-5-104 and IEC 61850 GOOSE, plus SWIFT MT and ISO 20022 message flows off the finance range.
PLC I/O and ladder state, setpoints, governor and excitation commands, breaker and relay positions, live process variables, sampled as the plant runs.
HMI and SCADA actions, alarm floods, acknowledgements, maker-checker steps: the human side of the loop, timestamped against the process.
Every red-team action mapped to MITRE ATT&CK for ICS: initial access, lateral movement, process manipulation, impact. Ground-truth by construction — we launched it.
Blue-team triage, the hypotheses, the containment decisions and the order they were taken in: the reasoning, not just the outcome.
Sensor and actuator readings from real PLCs and HO-scale plant: what moved, and what held fail-safe, when the attack landed.
What we ship
Prompt–response pairs grounded in real events, written to your taxonomy.
Full tool-use and decision traces: the defender loop as a sequence a model can imitate.
Expert-ranked good-vs-bad responses to the same incident state.
Held-out benchmark items with automated scorers and expert rubrics.
Aligned multimodal streams: packet, process variable, operator action, physical signal.
Every record carries its source scenario, labeller, and licence. No orphan data.
Benchmarks
The same engine that produces training data produces the held-out evaluation. Every task is drawn from a real campaign, scored by an automated grader plus an expert rubric, and sealed so it can't leak into the next crawl.
Spot the malicious setpoint change, the spoofed payment order, the mis-set interlock, hidden inside a stream of legitimate operations.
Given an unfolding incident, decide what to isolate, in what order, and what not to touch. Graded on decision quality, not vocabulary.
Read a pcap, an MT103, a GOOSE frame or a relay log and answer questions that only parsing it can settle.
Refuse or gate control actions that would trip a plant or move money unsafely — the OT equivalent of a harmlessness eval.
Diagnose the fault when the system has already collapsed fail-safe and the telemetry is partial and noisy.
Because these tasks can't be scraped, a benchmark you commission stays yours to measure against — release after release.
Co-design oneWhy us
Cyberange has spent a decade building the ranges the rest of the market sands away: physical plant, decoded industrial protocols, a standing red team. That infrastructure is the supply the engine runs on.
HO-scale ICS driven by real PLCs (the phygital range behind three Guinness World Records), plus live power, payments and rail simulators.
We launch the attack, so the label is known before the data is captured. No guessing what an event was after the fact, no noisy weak supervision.
Adversary emulation, DFIR and threat hunting on the same ranges. The defender reasoning in the data is expert reasoning.
Power-systems, financial-market-infrastructure and rail-signalling specialists label the data — not crowdworkers reading a rubric.
See the range for yourself: the Cyber Phygital Lab and the browser systems simulators are the same infrastructure this data comes off.
How we work
This is a partnership, not a download. We start from your capability gap and build the scenarios, the labels and the evaluation around it.
We map your capability gap to sector scenarios and concrete task definitions.
We run the campaigns on the range to your spec: new attacks, new operating conditions, the edge cases you care about.
To your schema and taxonomy: SFT, preference, agent-trajectory or eval, whichever you train on.
A training split plus a sealed benchmark, versioned and reproducible.
New scenarios as your model improves, with adversarial rounds that keep the benchmark off a plateau.
Provenance & safety
Every record is traceable to the scenario that produced it and ships with clean, transferable rights. No third-party scraping in the pipeline.
Data comes off synthetic ranges and lab plant. No real customer PII, no production credentials, nothing that maps back to a live network.
Benchmarks are sealed and never published to the open web, so a held-out eval stays held-out across model generations.
Scenario families can be built for a single partner and kept out of every other corpus we ship.
Questions
Tell us the capability you're trying to build. We'll scope the scenarios, generate the data, and hand back a benchmark you can measure against for years.