Cyberange · Data · For frontier models

Your model has never closed a breaker.

Frontier models learn cyber-physical operations from web text and generic logs. It's a world they've never touched. We run real attack-and-defend campaigns on live OT and critical-infrastructure ranges, then turn the signals into training data and benchmarks a model can learn from.

Power systems · Interbank payments & SWIFT · Rail signalling — more on the way

The gap

The data doesn't exist. So we make it.

Models are strong on code and prose because the internet is saturated with both. Cyber-physical operations are the opposite. The telemetry lives on air-gapped networks, incident data is confidential, and nobody publishes what a real ICS attack looks like from initial access to physical impact. So a model asked to reason about a tripping relay, a fraudulent MT103, or a mis-set signalling interlock is guessing by analogy — from text that describes these systems but never records one operating.

We close that gap at the source. Rather than scrape for events that were never recorded, we stage them on real plant, with expert operators on one side and a live adversary on the other.

Off-distribution

A substation control loop or a correspondent-banking settlement is nothing like the text a model was pretrained on.

Unlabelled where it exists

The rare real logs that surface have no ground-truth: which packet was the attack, which action was the fix.

Never end-to-end

No public source connects network, control, operator and physical layers through a single incident.

The engine

From live signals to LLM-ready corpora.

One pipeline turns an attack-and-defend campaign into data a model can train and be measured on. Every step is instrumented, so nothing is reconstructed after the fact.

  1. 01

    Run

    Red and blue teams operate real scenarios on our physical and virtual ranges: a substation, an interbank payment network, a metro line.

  2. 02

    Capture

    Every layer is captured at once: packet, protocol, process variable, operator action, physical sensor. All on the same clock.

  3. 03

    Label

    Domain experts annotate against attack ground-truth: technique, phase, intent, the correct response, the actual outcome.

  4. 04

    Transform

    Signals become model-ready: instruction pairs, agent trajectories, preference data, graded eval tasks, all in your schema.

  5. 05

    Deliver

    Versioned training corpora plus a sealed, held-out benchmark. Every record ships with its provenance and licence.

Signals in, formats out

Every layer of the incident, as data your stack accepts.

What we capture

Network & protocol

Full pcap plus decoded Modbus/TCP, DNP3, IEC 60870-5-104 and IEC 61850 GOOSE, plus SWIFT MT and ISO 20022 message flows off the finance range.

Process & control

PLC I/O and ladder state, setpoints, governor and excitation commands, breaker and relay positions, live process variables, sampled as the plant runs.

Operator telemetry

HMI and SCADA actions, alarm floods, acknowledgements, maker-checker steps: the human side of the loop, timestamped against the process.

Adversary trace

Every red-team action mapped to MITRE ATT&CK for ICS: initial access, lateral movement, process manipulation, impact. Ground-truth by construction — we launched it.

Defender trace

Blue-team triage, the hypotheses, the containment decisions and the order they were taken in: the reasoning, not just the outcome.

Physical ground truth

Sensor and actuator readings from real PLCs and HO-scale plant: what moved, and what held fail-safe, when the attack landed.

What we ship

Instruction / SFT pairs

Prompt–response pairs grounded in real events, written to your taxonomy.

Agent trajectories

Full tool-use and decision traces: the defender loop as a sequence a model can imitate.

Preference / RLHF sets

Expert-ranked good-vs-bad responses to the same incident state.

Eval tasks + graders

Held-out benchmark items with automated scorers and expert rubrics.

Event & time-series corpora

Aligned multimodal streams: packet, process variable, operator action, physical signal.

Schema & provenance

Every record carries its source scenario, labeller, and licence. No orphan data.

Benchmarks

Grade what a model would actually do.

The same engine that produces training data produces the held-out evaluation. Every task is drawn from a real campaign, scored by an automated grader plus an expert rubric, and sealed so it can't leak into the next crawl.

Attack detection in live telemetry

Spot the malicious setpoint change, the spoofed payment order, the mis-set interlock, hidden inside a stream of legitimate operations.

Incident-response reasoning

Given an unfolding incident, decide what to isolate, in what order, and what not to touch. Graded on decision quality, not vocabulary.

Protocol & log comprehension

Read a pcap, an MT103, a GOOSE frame or a relay log and answer questions that only parsing it can settle.

Safe actuation

Refuse or gate control actions that would trip a plant or move money unsafely — the OT equivalent of a harmlessness eval.

Root-cause under degraded mode

Diagnose the fault when the system has already collapsed fail-safe and the telemetry is partial and noisy.

Because these tasks can't be scraped, a benchmark you commission stays yours to measure against — release after release.

Co-design one

Why us

Ground-truth cyber-physical data can't be scraped. It has to be run.

Cyberange has spent a decade building the ranges the rest of the market sands away: physical plant, decoded industrial protocols, a standing red team. That infrastructure is the supply the engine runs on.

Real plant, not a diagram

HO-scale ICS driven by real PLCs (the phygital range behind three Guinness World Records), plus live power, payments and rail simulators.

We know the labels

We launch the attack, so the label is known before the data is captured. No guessing what an event was after the fact, no noisy weak supervision.

A decade of red and blue

Adversary emulation, DFIR and threat hunting on the same ranges. The defender reasoning in the data is expert reasoning.

Domain experts on staff

Power-systems, financial-market-infrastructure and rail-signalling specialists label the data — not crowdworkers reading a rubric.

See the range for yourself: the Cyber Phygital Lab and the browser systems simulators are the same infrastructure this data comes off.

How we work

A benchmark, built with you.

This is a partnership, not a download. We start from your capability gap and build the scenarios, the labels and the evaluation around it.

01

Scope

We map your capability gap to sector scenarios and concrete task definitions.

02

Generate

We run the campaigns on the range to your spec: new attacks, new operating conditions, the edge cases you care about.

03

Label & format

To your schema and taxonomy: SFT, preference, agent-trajectory or eval, whichever you train on.

04

Deliver & hold out

A training split plus a sealed benchmark, versioned and reproducible.

05

Refresh

New scenarios as your model improves, with adversarial rounds that keep the benchmark off a plateau.

Provenance & safety

Data you can put in a training run without a legal headache.

Provenance & licensing

Every record is traceable to the scenario that produced it and ships with clean, transferable rights. No third-party scraping in the pipeline.

Nothing routable to production

Data comes off synthetic ranges and lab plant. No real customer PII, no production credentials, nothing that maps back to a live network.

Contamination control

Benchmarks are sealed and never published to the open web, so a held-out eval stays held-out across model generations.

Exclusivity on request

Scenario families can be built for a single partner and kept out of every other corpus we ship.

Questions

Before you ask.

How is this different from buying raw logs or a public ICS dataset?
Public ICS datasets are small, stale, and unlabelled at the level a model needs. Raw logs are not training data. We generate fresh attack-and-defend campaigns on live ranges, capture every layer at once, and label against attack ground-truth we control, then ship the result as SFT, preference, agent-trajectory or evaluation data in your schema.
What sectors can you produce data for?
Today: electrical power systems, financial market infrastructure (interbank payments and SWIFT), and rail signalling and train control. New sectors are built as part of a co-design engagement using the same physical and virtual range infrastructure.
Is the data synthetic or real?
The signal is real, not model-generated: physical PLCs, decoded industrial protocols and live process variables, captured off a mix of real plant and high-fidelity simulators. None of it is synthetic text, and none of it contains production customer data.
Can we keep a benchmark private to us?
Yes. Scenario families can be built exclusively for one partner and withheld from every other corpus we ship, and benchmarks are sealed rather than published so they resist contamination across model generations.

Teach your model the systems that run the world.

Tell us the capability you're trying to build. We'll scope the scenarios, generate the data, and hand back a benchmark you can measure against for years.