# PROBE RESULTS: onnxruntime's phone-home, measured with nothing leaving the box

**What was measured:** whether an ordinary Python program using onnxruntime tries to reach Microsoft's
telemetry collector (`mobile.events.data.microsoft.com`), when it tries, how often, and whether
`ORT_DISABLE_TELEMETRY=1` stops it.

**Where and when:** the dev laptop, on 2026-09-28 (UTC, `date -u` at both ends).
- The probe ran from **08:45:25Z to 08:51:43Z**: preflight, then round 1 at 08:45:30Z, round 2 at 08:47:34Z,
  and round 3 at 08:49:39Z.
- The pre-registration was frozen at 08:45Z, before any onnxruntime process ran:
  `PREREG-probe.md`, sha256 `<withheld, see README.md>`. Every `result.json` records that same hash, and the file is
  unchanged since.
- This is the first measurement of its kind here. The recon (`RECON-onnxruntime-sources.md`) read the source
  and the local queue, but never ran onnxruntime.

## Verdict, one line per arm

- **1.30.0, default settings: it tries.** It looked up the collector and opened TLS to it 5 times in 120 s in
  every rep. The first try came 9.2 s after launch, 9.00 s after the import finished. **P1 CONFIRMED, 3/3.**
- **1.30.0 + `ORT_DISABLE_TELEMETRY=1`: silent.**
  - 0 DNS queries, 0 TLS/HTTP connections, and 0 network syscalls of any kind (strace, whole process tree).
  - It created no files: no deviceid, no queue DB, no `/tmp/.ses`.
  - **P2 CONFIRMED, 3/3.**
- **1.30.0 + `disable_telemetry_events()` right after import: it still tries.** 5 attempts per run, the first
  at 9.2 s, the same as default. Only the queued events change: the queue held just the import-time triple
  (ProcessInfo, RegisterEpLibraryStart/End), with no session events and no model canaries.
  **P3 CONFIRMED, 3/3.**
- **1.29.0, default settings: it tries**, identically to 1.30.0: 5 attempts, first at 9.2 s.
  **P4 CONFIRMED, 3/3.**
- **1.28.0, default settings (control): silent.** 0 queries, 0 connections, 0 network syscalls, and no files.
  **P5 CONFIRMED, 3/3.**
- **Instrument positive control (Python, not onnxruntime, connecting to the collector name on purpose): seen.**
  Every rep logged a lookup (A + AAAA) and a ClientHello with SNI `mobile.events.data.microsoft.com` at
  5.0 s. **P0 CONFIRMED, 3/3.**

All 18 self-tests passed, so no run is void. Every subject exited 0.

## The table, per arm and rep

Times are seconds after `t_launch`, the harness clock just before it started the subject.

- **Lookups** are queries the fake DNS server logged. Each attempt made two, A and AAAA, in the same
  millisecond.
- **Attempts** are TLS connections the fake collector accepted with SNI `mobile.events.data.microsoft.com`.
- **inet syscalls** are strace `connect()` calls to an AF_INET address. Every one was to `127.0.0.1:53` or
  `127.0.0.2:443`, never to any other address.

| Arm | Rep | Exit | Lookups | TLS attempts at (s) | inet syscalls | Files created (under the throwaway `/tmp`, which holds HOME) | Queue rows at exit |
|---|---|---|---|---|---|---|---|
| posctl | 1 | 0 | 2 | 5.0 | 2 | none | – |
| posctl | 2 | 0 | 2 | 5.0 | 2 | none | – |
| posctl | 3 | 0 | 2 | 5.0 | 2 | none | – |
| ort130-default | 1 | 0 | 10 | 9.2, 15.0, 24.5, 37.6, 77.4 | 15 | deviceid, onnxruntime.db, /tmp/.ses, /tmp/mat-debug-15.log | 12 |
| ort130-default | 2 | 0 | 10 | 9.2, 13.4, 22.8, 42.9, 80.1 | 15 | same four | 12 |
| ort130-default | 3 | 0 | 10 | 9.2, 13.6, 24.9, 47.4, 90.3 | 15 | same four | 12 |
| ort130-envoff | 1 | 0 | 0 | none | 0 | none | no DB |
| ort130-envoff | 2 | 0 | 0 | none | 0 | none | no DB |
| ort130-envoff | 3 | 0 | 0 | none | 0 | none | no DB |
| ort130-apioff | 1 | 0 | 10 | 9.2, 14.0, 23.4, 36.5, 79.4 | 15 | same four | 3 |
| ort130-apioff | 2 | 0 | 10 | 9.2, 12.7, 23.1, 43.9, 84.7 | 15 | same four | 3 |
| ort130-apioff | 3 | 0 | 10 | 9.2, 12.6, 21.3, 36.7, 83.6 | 15 | same four | 3 |
| ort129-default | 1 | 0 | 10 | 9.2, 14.2, 23.5, 42.1, 85.6 | 15 | same four | 12 |
| ort129-default | 2 | 0 | 10 | 9.2, 13.4, 24.4, 39.2, 76.9 | 15 | same four | 12 |
| ort129-default | 3 | 0 | 10 | 9.2, 12.7, 20.8, 33.4, 73.8 | 15 | same four | 12 |
| ort128-default | 1 | 0 | 0 | none | 0 | none | no DB |
| ort128-default | 2 | 0 | 0 | none | 0 | none | no DB |
| ort128-default | 3 | 0 | 0 | none | 0 | none | no DB |

- **Queue rows at exit:**
  - The 12 rows in the default arms were ProcessInfo, RegisterEpLibraryStart/End, ModelLoadStart/End,
    SessionCreationStart/SessionCreation/SessionCreationEnd, 2× EpDeviceUsage, RuntimePerf and SystemMetrics.
  - The 3 rows in the apioff arm were ProcessInfo and RegisterEpLibraryStart/End.
- **Machine-generated versions:**
  - `probe/runs/SUMMARY.md`: the same table plus the secondary table, regenerated by
    `python3 probe/summarize.py probe/runs`.
  - `probe/runs/run-all.log`: the live run's own summary.

**How often.** Every onnxruntime 1.29.0 and 1.30.0 run made exactly 5 attempts in its 120 s. (Only those two
versions were run; later versions are untested.) The gaps between them grew:

| Gap | Range across the 9 runs |
|---|---|
| 1st | 3.4 to 5.7 s |
| 2nd | 8.1 to 11.3 s |
| 3rd | 12.5 to 22.5 s |
| 4th | 37.2 to 46.9 s |

The last attempt fell between 73.8 and 90.3 s. The cadence past 120 s was not measured.

These are the ranges observed in the 9 runs, not bounds. The audit's 4 re-run reps (09:01Z to 09:12Z, same
kit) still made exactly 5 attempts each, growing gaps each time, but fell outside some of these ranges: 1st gap
3.05 s, 2nd 6.90 s, 3rd 22.83 s, 4th 33.82 s, and a last attempt at 69.13 s (`PROBE-AUDIT.md`, finding N-2).

## Secondary measures (pre-registered, not deciding)

- **S1, first attempt at 5 to 40 s: HELD, 9/9.**
  - The first attempt came 8.998 to 9.003 s after `import onnxruntime` returned, in all 9 runs (the SUMMARY
    table's "9.00 or 9.01" is display rounding).
  - The tighter anchor is the SDK's start, not the import's return: the first attempt came 9.010 to 9.014 s
    after `deviceid` was written (during the import), in all 9 runs and in the audit's 4 re-runs. In one audit
    re-run the import returned 72 ms after `deviceid`, and the attempt came 8.94 s after the import
    (`PROBE-AUDIT.md`, finding N-1).
  - **INFERENCE:** this matches the 1DS BEST_EFFORT timer's 9 s value for unmetered + charging (recon §2.1:
    `{36, 18, 9}`, upload at `timers[2]`). The container sees the host's sysfs, and the dev laptop's
    `/sys/class/power_supply/ADP1/online` read `1` at 08:46Z. Not tested on battery.
- **S2, more than one attempt in 120 s: HELD, 9/9.** Exactly 5 each.
- **S3, files: HELD.**
  - A1, A3 and A4 created these, in 9/9 runs:
    - `$HOME/.cache/Microsoft/DeveloperTools/.onnxruntime/deviceid`, at 0.14 to 0.22 s (during the import)
    - `onnxruntime.db`
    - `/tmp/.ses`, at the same moment as deviceid
  - A2 and A5 created nothing anywhere under `/tmp`, in 6/6 runs.
  - **Not predicted: an empty `/tmp/mat-debug-15.log`**, created at import in every A1/A3/A4 run and absent
    in A2/A5. 15 is the subject's PID inside the container; strace's final line for the main process is PID
    15. The lead's read of the Beat Lab voice's server found `/tmp/mat-debug-727.log` (0 bytes) for the Beat Lab's voice at PID 727, which
    fits the same pattern (`SWEEP-house.md` line 99; no such file for the restarted pid 3873263 that carries
    `ORT_DISABLE_TELEMETRY=1`). **INFERENCE:** the 1DS SDK ("MAT" is its older name) opens this debug log per
    process. Its source line is **UNSOURCED** here.
- **S4, queue after exit: HELD.**
  - A1 and A4 held session-level events in 6/6 runs. The `SessionCreation` payload carried all four canaries:
    - the graph name `probe_graph_canary`
    - the producer `probe-kit`
    - the metadata map entry `probe_meta_key=probe_meta_value_canary`
  - So the model's graph name, producer and custom metadata are what gets queued for upload. This is read from
    the local queue, not the wire.
  - A3 held only ProcessInfo and RegisterEpLibraryStart/End, with no canary, in 3/3 runs.
- **S5, raw identifiers absent: HELD.** No payload in any run contained the raw `deviceid` UUID, the run's
  raw `/etc/machine-id`, or the container hostname `ortprobe-host`.

**Descriptive, not pre-registered:**

- **What the queued payloads held:**
  - `"c:" + SHA-256(deviceid)` (uppercase) was in every row of every A1/A3/A4 run. That confirms the recon's
    derivation in 9/9 runs.
  - The `/tmp/.ses` UUID appeared as the SDK's installId in every such run's payload.
  - The interpreter path (`/probe/work/venvs/v1300/bin/python`) was in every row. On a host that path would
    carry the username (recon §2.4).
  - The ProcessInfo strings: `cpuModel` (the host's CPU string, via `/proc/cpuinfo`), `osDescription`
    (`Debian GNU/Linux 12 (bookworm)`, the container's), `architecture x86_64`, `deviceClass Desktop`,
    `Unmetered` / `Wired`, `runtimeVersion`, and SDK `EVT-Linux-C++-No-3.10.173.1`
    (`probe/runs/ort130-default-r1/queue/strings.txt`).
  - Field names **not** found in our SessionCreation payload (`/bin/grep -a -c` = 0 on the r1 DB copy):
    `modelGraphHash`, `modelWeightHash`, `modelFileName`, `modelDomain`, `modelProducerVersion` and
    `modelWeightType`. Our model was loaded from bytes and has no domain or version. The recon lists these from
    source. This probe neither confirms nor refutes them for file-loaded models.
- **Queue timing and retries:**
  - `retry_count` was 0 on every row at exit, despite 5 failed uploads per run.
  - The DB's mtime was 120.22 to 120.29 s, that is, written at teardown. That is consistent with recon §2.1
    (RAM first, persisted at flush/teardown).
- **Exit timing:** telemetry never delayed exit. The subject process ended 0.018 to 0.022 s after its own
  final timestamp in every onnxruntime arm, default included (posctl, which never imports onnxruntime:
  0.005 to 0.006 s). That is consistent with `CFG_INT_MAX_TEARDOWN_TIME = 0`.
- **The ClientHello:** 469 bytes every time, SNI `mobile.events.data.microsoft.com`, no ALPN, offering TLS 1.3
  and 1.2. Every attempt did a fresh A + AAAA lookup (no DNS reuse between attempts), on a separate thread.
  glibc's nscd probe (AF_UNIX `/var/run/nscd/socket`, ENOENT) appeared only with the first lookup.

## What this changes in the recon's open items

- Recon §6 said "Upload on our boxes is NOT proven either way". It is now **attempts proven, delivery not
  observed**.
  - The real PyPI wheels of 1.29.0 and 1.30.0, run on the dev laptop's kernel with default settings, open a TLS
    connection to the collector name within 9.2 s and keep retrying.
  - **INFERENCE:** any such process **on default settings** (no `ORT_DISABLE_TELEMETRY`) that lives past about
    9 s on a box with a route out makes the same attempts: for example the dev laptop's venvs, the long table's voice on
    its server (`SWEEP-house.md` H-1), and the Beat Lab's voice on its server **before** its 2026-09-28 08:00:24Z
    drop-in (`SWEEP-house.md` line 99). Whether the collector accepted them is outside what this probe can see.
- Recon §0.3 (import alone queues the triple) and §3 (`disable_telemetry_events()` leaves the uploader live and
  ProcessInfo still fires) are both **confirmed by measurement**:
  - the apioff arm attempted upload 5 times, and
  - its queue at exit held exactly that triple.
- On **1.30.0**, `ORT_DISABLE_TELEMETRY=1` set in the environment before the process starts is a **full**
  stop: no attempts, no identifiers, no files (3/3, pre-registered). Setting it from Python before the import
  was not a separate arm.
  - **1.29.0 with the variable set was not a pre-registered arm.** The Beat Lab's voice and the long table's voice run
    1.29.0. The audit ran it once, as an exploratory reading outside the pre-registration: silent. It made 0
    DNS queries, 0 connections and 0 network syscalls, created no files, and its network namespace's kernel
    counters did not move from launch to exit. A 1.29.0 default run beside it made 5 attempts
    (`PROBE-AUDIT.md`, finding I-1). That is n = 1, so treat it as a strong hint, not a verdict.

## Exact environment

| Item | Value | Source |
|---|---|---|
| Host | the dev laptop (hostname withheld), Intel Core Ultra 9 290HX Plus, 24 CPUs, MemTotal 64203440 kB, AC adapter online | `nproc`, `/proc/cpuinfo`, `/proc/meminfo`, `/sys/class/power_supply/ADP1/online`, 08:46Z |
| Host OS | Ubuntu 26.04.1 LTS | `/etc/os-release` |
| Kernel | `Linux 7.0.0-34-generic #34-Ubuntu SMP PREEMPT_DYNAMIC Wed Sep 2 14:29:37 UTC 2026 x86_64` | `uname -srvm`; the containers report the same release |
| Docker | client and server 29.1.3 (rootful daemon; the user is in the `docker` group) | `docker version` |
| Container image | `debian:12-slim`, `sha256:88200866dfff7ea7f5cbcb6ec7c8a701889efe6fe859fe64d6990e4b07ea4171` | `docker image inspect` |
| Container flags | `--network none --user 1000:1000 --cap-drop ALL --security-opt no-new-privileges --read-only --tmpfs /tmp --memory 1g --cpus 2 --pids-limit 512 --hostname ortprobe-host`, with a fresh random `/etc/machine-id` per run | `probe/run-arm.sh` |
| Python | CPython 3.12.13 (python-build-standalone, the one uv manages on this host), mounted read-only at `/probe/python` | `result.json` `harness_python` |
| strace | 6.19, the host binary run via its own loader inside the container | `strace -V` |
| Packages (identical in all 3 venvs except onnxruntime) | flatbuffers 25.12.19, ml_dtypes 0.6.0, numpy 2.5.3, onnx 1.23.0, packaging 26.3, pip 25.0.1, protobuf 7.36.2, typing_extensions 4.16.0 | `result.json` `packages` |
| onnxruntime wheels (PyPI) | 1.30.0 `fa688e78…9328` (manylinux_2_28), 1.29.0 `2b80d8c7…acb`, 1.28.0 `0a83bdb7…933c` | `probe/runs/wheels-SHA256SUMS` (full hashes) |
| `onnxruntime_pybind11_state.cpython-312-x86_64-linux-gnu.so` | 1.30.0 `0b2a6e0d…b33b`, 1.29.0 `ff54b93f…90ee`, 1.28.0 `6a31ea84…fd83`; the collector string occurs 3/3/0 times, `ORT_DISABLE_TELEMETRY` 1/1/0, `DO_NOT_TRACK` 0/0/0 | `probe/runs/ort-pybind-SHA256SUMS`; `/bin/grep -a -c -F` at 08:35Z |

**Install path.** Wheels were `pip download`ed on the host, with the network, at about 08:43Z. The venvs were
then built **inside a `--network none` container** with `pip --no-index`, so no install step could touch
the network either.

## The instrument, and one deviation from the lead's design

The details are in `PREREG-probe.md`, "Instrument" and "Amendment 0".

- The lead asked for `unshare --user --map-root-user --net --mount`. On the dev laptop it fails:
  - `kernel.apparmor_restrict_unprivileged_userns=1` → `unshare: write failed /proc/self/uid_map: Operation
    not permitted`.
  - bwrap works, but its AppArmor profile strips capabilities from its children. `ip_unprivileged_port_start`
    stays 1024, so binding 53 or 443 gives EACCES.
- I did not use `nsenter` into bwrap's user namespace to get capabilities back. That would deliberately
  sidestep the AppArmor restriction.
- The working form is **Docker `--network none`**, where the container netns defaults to
  `ip_unprivileged_port_start=0`. Isolation is the same kernel mechanism: one interface (`lo`) and zero routes.
  It is checked in every run.
- The design gained **strace**, which the lead's design did not have. It sees connects to any address,
  including hard-coded IPs that DNS logging would miss.

**No-egress evidence:**

- `probe/runs/preflight.txt`, 08:45:25Z: a real curl in a `--network none` container.
  - `https://1.1.1.1/` → exit 7, "Could not connect to server".
  - `https://20.184.175.9/` → exit 7.
  - `https://mobile.events.data.microsoft.com/` → exit 28, "Resolving timed out".
- **Per-run self-test** (in each `result.json` `selftest`, 18/18 pass):
  - TCP to `1.1.1.1`, `20.184.175.9` and `20.184.175.13` on 443 → `errno 101 Network is unreachable`.
  - Interfaces `["lo"]`, 0 IPv4 routes, 0 non-lo IPv6 routes.
  - The fake DNS answered `127.0.0.2` and logged the query.
  - The fake TLS listener logged the self-test SNI.
- In every run, strace saw no AF_INET/AF_INET6 destination other than `127.0.0.1:53` and `127.0.0.2:443`.
- The dev laptop's own ORT queue (`~/.cache/Microsoft/DeveloperTools/.onnxruntime/onnxruntime.db`) still has mtime
  2026-09-28 01:00:47Z. The probe never touched it: each container had its own tmpfs HOME.

## What this probe does NOT show

- **Attempts, not delivery.** The listener hung up mid-handshake, so we cannot say what the real collector
  would have accepted, what it would have answered, or what a completed TLS session carries on the wire.
- **The payload is read from the local queue DB after exit, not from the wire.** The queue holds what the SDK
  meant to send. Whether one upload batches all of it is not observed.
- **A container is not a host:**
  - `/etc/os-release` is Debian 12, so ProcessInfo said Debian.
  - The machine-id is a fresh fake, HOME is empty, `/run` is absent, and there is no host DNS stack.
  - `/proc/cpuinfo`, `/proc/meminfo` and `/sys` are the host's, so the CPU string, RAM and power state are
    real.
  - Network type and cost read `Wired` / `Unmetered` even with only `lo`. **INFERENCE:** the SDK's Linux
    network info does not reflect the interface. **UNSOURCED** in code.
- **120 s per run.** Not measured:
  - the cadence past 90 s
  - the 10-minute RuntimePerf cap
  - the SDK's drop after more than 5 failed retries (our rows all read `retry_count` 0 at exit)
  - behaviour on battery (the 18 s timer)
  - an import-only process with no session
- **Every run was a first run** (a fresh HOME each time): deviceid Status is New, and there were no leftover
  queued events from an earlier process. The recon's "sent on the next run" path was not exercised.
- **Only the CPU wheels, on Python 3.12, on x86_64.** Not tested: onnxruntime-gpu, onnxruntime-genai,
  macOS, Windows, and setting the variable via `os.environ` inside Python before the import.
- **Six containers ran at once per round.** They could not see each other, since each had its own netns, but
  timings are under that mild CPU contention.

## Reproduce

Needs Docker, the `debian:12-slim` image, a host CPython 3.12 with pip, strace, and, for the preflight, the
`curlimages/curl` image.

```
cd probe
PROBE_WORK=/some/scratch PROBE_PYTHON_DIR=/path/to/cpython-3.12 ./prepare.sh   # network used here only
PROBE_WORK=/some/scratch PROBE_PYTHON_DIR=/path/to/cpython-3.12 ./run-all.sh 3 # preflight + 3 rounds × 6 arms, ~7 min
python3 summarize.py runs
```

Scratch here was `<a scratch directory>`, about 681 MB. It was **deleted** at 08:53Z with
the wheels, venvs and strace bundle, and no probe container remains (`docker ps -a`: 0).

## Files

- `PREREG-probe.md`: the pre-registration, frozen at 08:45Z, sha256 `<withheld, see README.md>`.
- `probe/`: the kit.
  - `prepare.sh`: wheels, strace bundle, offline venvs.
  - `preflight.sh`
  - `run-arm.sh`: one arm, one rep.
  - `run-all.sh`
  - `inside.py`: the listeners, self-test, strace launch and snapshots.
  - `subject.py`: the workload.
  - `summarize.py`
  - `resolv.conf`
- `probe/runs/<arm>-r<rep>/`, one folder per run:
  - `result.json`: everything, including the self-test.
  - `strace.log`
  - `marks.json`: the subject's own timestamps.
  - `harness.log`
  - `subject.stdout` and `subject.stderr`
  - `machine-id`: the fake one.
  - For telemetry arms only:
    - `queue/onnxruntime.db`: a copy of the queue DB.
    - `queue/strings.txt`
    - `deviceid` and `tmp.ses`: throwaway identifiers.
- `probe/runs/SUMMARY.md`, `run-all.log`, `preflight.txt`, `wheels-SHA256SUMS`, `ort-pybind-SHA256SUMS`, and
  `dryrun/`: the instrument check at 08:44Z, before the freeze, with no onnxruntime.

**Kit changes after the runs** (both after round 3 ended at 08:51:43Z, from the lane's own transcript):
- 08:52:14Z, `summarize.py`: a `secondary()` table (S1 to S5 plus descriptive columns) was appended. The
  verdict code (the `EXPECT_*` maps, the `ok` tests and `N_SECONDS`) was not touched, so `run-all.log`'s live
  verdicts and `SUMMARY.md`'s agree.
- 08:53:08Z, `inside.py`: the `container_os` parse now strips the trailing quote and newline, which the
  recorded results still show as `Debian GNU/Linux 12 (bookworm)"\n`. It is cosmetic and affects no
  measurement.

## Before any public use

- `probe/runs/preflight.txt`, `probe/runs/dryrun/preflight.txt` and `probe/runs/run-all.log` name the host
  (`<hostname withheld>`). Redact that.
- The queue copies and `strings.txt` carry:
  - the host's CPU model string
  - the in-repo public tenant token `o:5ad963bd…`
  - throwaway container identifiers only: no username and none of the host's ids
- The kit scripts carry no personal paths (`/bin/grep -a -r -i <an operator's account name> probe/*.sh probe/*.py` → 0).
- The audit's `audit/witness/audit-sandbox.json` names the dev laptop's LAN and private-network addresses (they were
  connect targets in the no-route check). Redact them, or leave `audit/` out.

## Audit addendum, 2026-09-28 (independent audit, 08:56Z to 09:16Z)

**What was checked:** an independent audit re-checked this probe's no-egress sandbox, its pre-registration
timing, its instrument, and every figure above against the raw run files. It also re-ran 1.30.0 default,
env-off, api-off and the positive control. The full report is `PROBE-AUDIT.md`, and the evidence is in `audit/`.

**Verdict: HOLDS.** All six pre-registered predictions stand. The re-runs behaved like the original 9 runs:
5 attempts with the first at 9.2 s for default and api-off, and silence for env-off. A kernel-counter
witness read from the host found zero packets and zero connection attempts from the env-off arm's network
namespace, from launch to exit.

**What changed in this file:**
- The `ORT_DISABLE_TELEMETRY` "full stop" line is now scoped to 1.30.0. 1.29.0 with the variable set gets
  one exploratory, not-pre-registered audit reading, which was silent.
- The first-attempt anchor is now the SDK start (+9.010 to 9.014 s), not the import's return.
- The gap ranges are now marked as observed ranges, not bounds.
- The exit-timing and "≥1.29" wording is now scoped.
- The Beat Lab voice inference is scoped to default settings.
- The `summarize.py` post-run edit is now disclosed.
- The `mat-debug-727` observation now cites its source.
