# ONNX Runtime's telemetry on Linux, measured — data kit

**What this directory holds.** The sealed test behind the exhibit *ONNX Runtime's
telemetry on Linux, measured: on by default since 1.29, and the switch that stops
it*, whole: the seven scripts and the resolver file that ran it, the
pre-registration frozen before the first run, every text file of all 18 runs and of
the dry run before them, the summaries, the measurement record, and the independent
audit with its own scripts, its 8 re-runs and its kernel-counter witness.

The test asked one question of the official PyPI packages of `onnxruntime` 1.28.0,
1.29.0 and 1.30.0 for Linux x86-64, on CPython 3.12.13: does a default install try
to reach Microsoft's telemetry collector, `mobile.events.data.microsoft.com`, and
does `ORT_DISABLE_TELEMETRY=1` stop it? Every run was a Docker container with no
network at all, beside a stand-in name server and a stand-in collector that logged
every attempt, read its first handshake message, and hung up before any encrypted
session began.

Everything here happened on 2026-09-28, in UTC:

| What | Opened (UTC) | Closed (UTC) |
|---|---|---|
| The instrument dry run, no onnxruntime | 08:44Z | 08:44Z |
| The pre-registration frozen | 08:45Z | 08:45Z |
| The probe: preflight, then three rounds of six rows | 08:45:25Z | 08:51:43Z |
| The independent audit | 08:56Z | 09:16Z |
| The audit's re-runs, inside that window | 09:01Z | 09:12Z |

**Licence: CC BY 4.0.** Take these files, re-run the test, check our arithmetic,
publish what you find. Attribution: strata→signal research,
research.strata2signal.com. If you find an error, we want to hear about it:
hello@strata2signal.com.

Machine readers start at `index.json`, which lists every file with its byte count
and its sha256 — every file but itself, `provenance.json` and the browsable
`index.html`, which its `file_count` counts all the same. `provenance.json` is the
other half: it names every rule applied to every file, with the count of times each
fired, and says of every file whether a rule touched it.

## Fetch the kit

The files are served one by one beside this README. This fetches all of them into
the current directory and checks every sha256 against `index.json` as it goes; set
`KIT` to this directory's own address, ending in `data/`:

```
export KIT=https://research.strata2signal.com/<the page>/data/
python3 -c 'import hashlib, json, os, urllib.request as u
b = os.environ["KIT"]; m = json.load(u.urlopen(b + "index.json"))
for f in m["files"]:
    d = u.urlopen(b + f["file"]).read()
    assert hashlib.sha256(d).hexdigest() == f["sha256"], f["file"]
    os.makedirs(os.path.dirname(f["file"]) or ".", exist_ok=True)
    open(f["file"], "wb").write(d)
print(len(m["files"]), "files fetched, every sha256 checked")'
```

A download carries no file modes, so make the scripts executable before you run
them: `chmod +x probe/*.sh probe/*.py`.

## Run the sealed test yourself

You need x86-64 Linux with Docker (ours was rootful Docker 29.1.3, the account in
the `docker` group); the `debian:12-slim` and `curlimages/curl` images; `strace` on
the host, which `prepare.sh` copies into the containers with its libraries; and a
**relocatable** CPython 3.12 with pip on the host, because `debian:12-slim` has no
Python of its own and the scripts mount yours read-only into every container. Ours
was CPython 3.12.13 from python-build-standalone, the build `uv` installs. Our
scratch reached about 681 MB (`PROBE-RESULTS.md`, *Reproduce*).

```
cd probe
chmod +x *.sh *.py
export PROBE_WORK=/some/scratch PROBE_PYTHON_DIR=/path/to/cpython-3.12
./prepare.sh                                  # the only step that uses the network
PROBE_RUNS=/some/scratch/runs ./run-all.sh 3  # preflight, then 3 rounds of 6 rows
```

Three rounds take about seven minutes, six containers at a time. Then
`python3 summarize.py /some/scratch/runs` prints the table, and
`python3 ../audit/verify.py /some/scratch/runs` re-derives it from the raw files.

**Point `PROBE_RUNS` at an empty directory.** It defaults to `probe/runs/`, which
holds our 18 runs, and forgetting it fails quietly: `preflight.sh` writes its record
over our `probe/runs/preflight.txt`, every `run-arm.sh` refuses to write into a run
directory that is not empty and says so only on standard error, and `run-all.sh`
still ends by reprinting OUR table as if it were yours. The *Reproduce* block in
`PROBE-RESULTS.md` predates this kit and leaves `PROBE_RUNS` out; add it.

**CI variables.** The library switches its own telemetry off when it sees `CI`,
`GITHUB_ACTIONS` or eleven other such variables set (the exhibit page gives the
rule), so a program run inside CI can look silent whatever the switch says.
`run-arm.sh` passes no variable into the container, and `inside.py` hands the test
program an environment built from nothing (`PATH`, `HOME`, `LANG`,
`PYTHONDONTWRITEBYTECODE` and the row's own variable), so no CI variable reaches it;
if you run `subject.py` any other way, run it outside CI.

Your runs record the sha256 of this directory's `PREREG-probe.md`. Ours recorded the
sha256 of the original, which this kit withholds; the section after next says why.

## Check our own runs without running anything

```
python3 probe/summarize.py probe/runs          # reprints probe/runs/SUMMARY.md byte for byte
python3 probe/as-run/summarize.py probe/runs   # reprints the table at the foot of probe/runs/run-all.log
python3 audit/verify.py probe/runs             # reprints audit/verify-orig.txt, less its queue lines
```

The first prints 7,883 bytes identical to `probe/runs/SUMMARY.md`, and the second
4,883 bytes identical to the table `run-all.sh` printed live at 08:51Z; both were
re-checked byte for byte over the copies in this kit. `verify.py` prints its two
queue lines per run only where a copy of that run's queue database is present, and
this kit does not carry those copies (below); every other line it prints over these
files, and over `audit/runs` against `audit/verify-audit.txt`, is identical to what
the audit recorded.

## The pre-registration's fingerprint, withheld

*Added 2026-09-28 (UTC), before the page first linked this kit.* This section is
about one value: the sha256 of the pre-registration that every run wrote into its
`result.json` as `prereg_sha256`, so that an edit made after the first run would
show.

**In this kit that value is withheld.** The field reads `null` in all 18 of the
probe's runs and in the six audit re-runs that recorded it, and the measurement
record and the audit quote it as `<withheld, see README.md>`. The audit's two 1.29.0
runs came from a scratch copy of `run-arm.sh` and recorded `none`, which stands. The
dry run's value is withheld too: it is the sha256 of the draft before the freeze,
which this kit does not carry. An earlier cut of this kit, made the same day and
never linked, printed the recorded value in this section; this cut does not.

**Why.** The recorded value is the sha256 of the private original, and the copy here
differs from that original only in three words, which the rules below replaced: the
name of the machine it was written on, twice, and an operator's account name, once.
With the original's sha256 beside a copy like that, anyone could put guesses back
in, hash, and compare, offline, until one matched. It is the same reason no file
here carries the digest of any original a rule changed (below).

**What it costs you.** You cannot check the order of events from this kit alone. The
proof of order is kept in the private record: the recorded sha256, the same in all
18 runs and in the audit's six re-runs up to 09:05Z, and the file's own times, its
last write at 08:45:16.67Z against the first test program's launch at 08:45:31.51Z.
The independent audit checked it there, and its table of those times is published
in full in `PROBE-AUDIT.md`, check 2. The copy of `PREREG-probe.md` here is that
frozen text with the three words replaced and nothing else moved: `provenance.json`
gives the rules that fired in it and their counts, and `index.json` gives the copy's
own sha256.

## The two scripts edited after the runs

Both edits came after round 3 ended at 08:51:43Z, and neither touched the code that
decides a verdict. The bytes that ran are here too, in `probe/as-run/`, recovered by
replaying the probe agent's own recorded file writes; applying the one recorded edit
to each gives the copy in `probe/` byte for byte.

| Script | As it ran: sha256, bytes | Edited (UTC) | What the edit did | As published: sha256, bytes |
|---|---|---|---|---|
| `probe/summarize.py` | `e864ea6e5b23c6fd46d99968b23170756fb49ae965225e0dfda7e30c8dec3776`, 4,887 | 08:52:14Z | Added the secondary table and nothing else: 41 lines inserted after line 93, one call at the end of `main()` and a new `secondary()` function. No existing line changed, the verdict code included. | `20300506248af87bd135a80cb56c6915c1f338c87bd60ec9eff03fbd4321b52c`, 7,446 |
| `probe/inside.py` | `b67df717833dbb78242f602110f8f60e64eb6584f9c8d012a56abe2937ad4ba4`, 18,954 | 08:53:08Z | One line: `container_os` now strips the trailing newline before the quote. The 18 runs still show the old form, `Debian GNU/Linux 12 (bookworm)"\n`; the audit's 8 runs, made after it, show the new one. | `346aeef9843fae720ca2022eab6edf86a161f16b0d465c75d854a2fe66babfa4`, 18,962 |

The edit times are the moments the agent issued each edit; the files' own modified
times read 08:52:18.92Z and 08:53:11.48Z. The other five scripts and the resolver
file are the bytes that ran.

## What is in here

| Path | What it is |
|---|---|
| `PREREG-probe.md` | The pre-registration: the question, the instrument, the six rows, the six predictions, the void rules, and Amendment 0 (why Docker and not `unshare`). |
| `PROBE-RESULTS.md` | The measurement record, written by the agent that ran the probe and corrected by the audit: the verdicts, the table per run, the gaps between attempts, the secondary measures, the exact environment, and what the probe does not show. |
| `PROBE-AUDIT.md` | The independent audit: the sandbox re-checked from inside and from the host, the pre-registration's timing, three positive controls, the re-runs, every figure matched against the raw files, and the kernel-counter witness. |
| `probe/prepare.sh`, `preflight.sh`, `run-arm.sh`, `run-all.sh` | Fetch and build (the only network use), prove the container reaches nothing, run one row, run them all. |
| `probe/inside.py` | The in-container harness: the stand-in name server and TLS and HTTP listeners, the self-test that voids a run on any failure, the test program under `strace`, and the after-run reading of files and the queue. |
| `probe/subject.py` | The test program: imports onnxruntime, builds a one-addition model in memory carrying four planted marker strings in three places (the graph name, the producer name, and one metadata key and its value), runs it once, stays alive for 120 s. |
| `probe/summarize.py` | The summariser: a row per run, the verdict per prediction, the secondary table. |
| `probe/resolv.conf` | Mounted over each container's `/etc/resolv.conf`: the only resolver is the stand-in. |
| `probe/as-run/` | `inside.py` and `summarize.py` exactly as they ran (the section above). |
| `probe/runs/<row>-r<n>/` | One run each, 18 in all: `result.json` (everything, the self-test included; its `prereg_sha256` is withheld, above), `strace.log`, `marks.json` (the test program's own timestamps), `harness.log`, the empty `subject.stdout` and `subject.stderr`, `machine-id`, and for the rows that ran telemetry `deviceid`, `tmp.ses` and `queue/strings.txt`. |
| `probe/runs/SUMMARY.md`, `run-all.log`, `preflight.txt` | The summariser's output, the live log, and the preflight's record. |
| `probe/runs/wheels-SHA256SUMS`, `ort-pybind-SHA256SUMS` | The sha256 of every wheel downloaded and of each version's compiled library. |
| `probe/runs/dryrun/` | The instrument dry run at 08:44Z, before the freeze, with no onnxruntime. |
| `audit/audit_inside.py`, `audit/verify.py` | The audit's no-route and positive-control checks, and its re-derivation of every figure from the raw files. |
| `audit/verify-orig.txt`, `audit/verify-audit.txt` | `verify.py`'s recorded output over the 18 runs and over the audit's 8. |
| `audit/runs/` | The audit's 8 runs, the same files per run as above, and each launch's one-line result. |
| `audit/witness/` | `audit-sandbox.json` (the no-route check and the positive controls) with its empty error stream, and four logs of a container's own kernel network counters, read from the host about once a second. |
| `audit/run-arm-audit-arm.diff`, `audit/*-SHA256SUMS` | The one line the audit added for its exploratory 1.29.0 run with the switch set, and the sha256 of its own fresh downloads, which match the probe's. |

## What was replaced, and why

This kit ships only once our machine names and addresses are out of it. Every file
was copied from the probe's and the audit's own record and rewritten only by named,
mechanical rules, applied by one program (the hub's `tools/kit_sanitise.py`) in a
fixed order. `provenance.json` names each rule, says what it does, and counts its
firings in every file. The literal each rule replaces is not published, and neither
is the digest of any file a rule touched: a digest of a redacted original would let
anyone test guesses for what was replaced. The same holds for a digest the record
itself carries, so the sha256 the runs recorded of the pre-registration is withheld
too (above).

| What was replaced | Became |
|---|---|
| The dev laptop's host name (6 times, in the preflight's first line and the records quoting it) | `<hostname withheld>` |
| The estate names of the dev laptop and three other machines (23) | The role each played: the dev laptop, the inference box, and the servers behind the Beat Lab's voice and the long table's voice |
| An operator's account name (2) | An operator's account, or `<an operator's account name>` inside a quoted command |
| The name of the private network the dev laptop is on (3) | Private network |
| Three of the dev laptop's own addresses, dialled by the audit's no-route check (4) | `<the dev laptop's LAN address>`, `<the dev laptop's private-network address>`, `<the dev laptop's second Docker bridge address>` |
| Two services' unit names (11) | What the page calls them: the Beat Lab's voice, the long table's voice |
| An agent transcript's file name (1) and a scratch path (1) | The probe agent's transcript; `<a scratch directory>` |
| The sha256 the runs recorded of the pre-registration (27: whole in 24 `result.json` files and in the audit, shortened twice in the measurement record), and the one the dry run recorded of the draft before the freeze (2) | `null` in `result.json`, and `<withheld, see README.md>` in prose |
| The throwaway identifiers made for each sealed container: 27 machine IDs, 13 device IDs with their 120 uploaded hashes and 13 hash prefixes, and 13 install IDs in 146 places | A placeholder naming the run, such as `<deviceid:ort130-default-r1>`, the same everywhere that ID occurred |

The container identifiers were random and made fresh for each run — the machine ID
by `run-arm.sh` from the host's random source, the device and install IDs by the
library inside the container — for a container with no network that was then
deleted, and none reached anyone. They are replaced anyway, and consistently, so a
file that repeated an ID still repeats its placeholder (the summariser's install-ID
column reads "yes" on these files exactly as on the originals). A run of your own
mints fresh ones, and there you can check that the queued payloads carry `c:` and
the uppercase SHA-256 of the device ID, as the audit did in every run. The queued
events' own random IDs — one per event, plus two per process — are left as written:
each occurs in one run only, and none came from the dev laptop. No identifier of
the dev laptop itself is left in these files: not its own device ID, machine ID or
`/tmp/.ses`, not its network cards' hardware addresses, and no address of its
interfaces other than Docker's default `docker0` address (below). That was checked
against the live machine before the cut.

`PROBE-RESULTS.md` ends with a list headed *Before any public use*: the host name to
redact, the addresses in `audit/witness/audit-sandbox.json`. This kit's rules are
what carried that list out, so its items read as done here.

**No figure moved.** The program's fence reads every number in every file on both
sides and refuses on any difference outside the replaced spans; over the 266 files
the rules produced it read 51,185 numbers, all identical. The only digits any rule
removes are inside the names, addresses, identifiers and withheld sha256 values
above. The other
three files here are this README and the two recovered scripts in `probe/as-run/`,
which no rule touched.

**Left as written, on purpose:** Docker's default `docker0` bridge address, which
most Docker hosts keep and which the audit's script dials on whatever host runs it;
the container's throwaway home folder, `/tmp/home`; the collector's two addresses
and the public resolvers the self-test dialled; the dev laptop's processor model,
memory size, kernel build and Docker version, which are hardware and software
versions and not a name or an address; and the names of two private working
documents the records cite, `RECON-onnxruntime-sources.md` (this workshop's reading
of the library's source, which the page's Sources section states in public) and
`SWEEP-house.md` (a survey of this workshop's own machines).

## What is not here

- **The queue database copies**, `queue/onnxruntime.db` in 13 runs. They are SQLite
  files, and every file here passed through text rules and the site's redaction
  gate, which reads text. What each queued payload carried is in
  `queue/strings.txt`. Each row's facts — event names, retry counts, the marker
  strings, and whether it carried the device-ID hash, the interpreter path, the raw
  device ID, the machine ID or the host name — are in that run's `result.json`;
  whether it carried the install ID is read from `queue/strings.txt` and the run's
  `tmp.ses`, which is what the summariser does.
- **The wheels, virtual environments and strace bundle.** `prepare.sh` fetches the
  same wheels from PyPI, and `probe/runs/wheels-SHA256SUMS` gives each one's sha256.
- **The working documents and the agents' transcripts** the records cite.

## The rules for reading a figure

- A **lookup** is a query the stand-in name server logged; every attempt came with
  two, A and AAAA.
- An **attempt** is a TLS connection the stand-in collector accepted with the server
  name `mobile.events.data.microsoft.com`. It read the first handshake message and
  hung up before any encrypted session began, so an attempt is never a delivery.
- A **network system call** is any `connect`, `sendto`, `sendmsg` or `sendmmsg` by
  the test program's whole process tree to an IPv4 or IPv6 address, from `strace`.
  Local AF_UNIX calls are counted apart.
- **Times** are seconds after `t_launch`, the harness's clock just before it started
  the test program.
- A **prediction** holds only if it holds in every run of its row.

## The limits of this run

One library at three versions, on one laptop's processor on mains power, in a
Debian 12 container with no network, 120 seconds a run of onnxruntime (the
instrument check held 25 s), every run a first run with an empty home folder. It
shows attempts, not what the collector would have accepted, and the queue's
contents as the library stored them, not as they would have crossed the wire.
`PROBE-RESULTS.md`'s own section, *What this probe does NOT show*, lists the rest.
