# PROBE AUDIT: the no-egress probe of onnxruntime's phone-home

**What was audited:** the probe in `PREREG-probe.md` and `PROBE-RESULTS.md`. It measured whether onnxruntime
1.29.0 and 1.30.0 try to reach Microsoft's telemetry collector (`mobile.events.data.microsoft.com`), and
whether `ORT_DISABLE_TELEMETRY=1` stops them. Its runs went from 08:45:25Z to 08:51:43Z.

**When and where:** the dev laptop, on 2026-09-28 from 08:56Z to 09:16Z (UTC, `date -u`). This was an
independent audit lane. It did not touch the inference box or any other box, used no sudo, and changed no configs.

## Verdict: HOLDS

- All six pre-registered predictions stand.
- The sandbox had no route out, as re-verified from inside and from the host.
- The pre-registration was frozen 15 s before the first subject launched.
- The instrument sees an attempt through each of three paths: DNS + SNI, HTTP Host, and strace for a
  hard-coded IP.
- The audit's own re-runs reproduced the arms.
- Every figure in `PROBE-RESULTS.md` matched the raw files.

There are no BLOCKING findings. There are 2 IMPORTANT findings, both about the scope of the claims, and 10
NICE ones. Every fix the audit made itself is mechanical, and each is marked FIXED below.

## Findings

| Id | Severity | Finding | Status |
|---|---|---|---|
| I-1 | IMPORTANT | The env switch was measured on 1.30.0 only, but the "full stop" line and the build return's headline did not name a version. The Beat Lab's voice and the long table's voice run 1.29.0 (`SWEEP-house.md` lines 99, 117). | FIXED: scoped to 1.30.0. One exploratory, not-pre-registered 1.29.0 env-off reading was added: silent. |
| I-2 | IMPORTANT | The INFERENCE line said the Beat Lab's voice on its server "makes the same attempts". The Beat Lab's voice has carried `ORT_DISABLE_TELEMETRY=1` since 08:00:24Z (`SWEEP-house.md` line 99). | FIXED: scoped to default settings, and to the Beat Lab's voice before its drop-in. |
| N-1 | NICE | The first attempt tracks the SDK's start, not the import's return. It came 9.010 to 9.014 s after `deviceid` was written in 13/13 runs, but 8.94 to 9.003 s after the import returned. "9.00 or 9.01" is display rounding. | FIXED |
| N-2 | NICE | The gap ranges and the 73.8 to 90.3 s last-attempt range are ranges observed in 9 runs, not bounds. The audit's re-runs fell outside them (1st gap 3.05, 2nd 6.90, 3rd 22.83, 4th 33.82; last attempt 69.13 s). | FIXED: labelled as observed ranges |
| N-3 | NICE | "0.018 to 0.022 s … in every arm": posctl reads 0.005 to 0.006 s. | FIXED: "every onnxruntime arm" |
| N-4 | NICE | "Every onnxruntime ≥1.29 run" implies versions that were never run. | FIXED: "1.29.0 and 1.30.0" |
| N-5 | NICE | `summarize.py` was edited after the runs (08:52:14Z) and this was not disclosed. The edit appended the secondary table only. The verdict code was not touched (lane transcript). | FIXED: disclosed |
| N-6 | NICE | The `mat-debug-727.log` observation carried no source. | FIXED: cites `SWEEP-house.md` line 99 |
| N-7 | NICE | The recorded `container_os` ends in a quote **and** a newline (`…(bookworm)"\n`). The results named only the quote. | FIXED |
| N-8 | NICE | The pre-registration says the port-start read was "verified 08:36Z". The transcript shows 08:35:16Z. The ".so grep at 08:35Z" ran on uv-installed copies that were deleted at 08:43:27Z. Their sha256 equals the probe's pip-built copies, and the audit's re-grep of fresh PyPI copies matches (1.30.0 and 1.29.0: collector 3, `ORT_DISABLE_TELEMETRY` 1, `DO_NOT_TRACK` 0). | No fix: the pre-registration is frozen, and the claims are consistent |
| N-9 | NICE | The dev laptop's own queue: `.db` 01:00:47Z matches. `-wal` (07:54:37Z) and `-shm` (07:54:52Z) also exist. Both predate the probe, so "untouched by the probe" holds. What wrote them at 07:54Z is outside this audit. | Info |
| N-10 | NICE | strace `%network` cannot see io_uring sockets, and the original runs had no witness below the syscall layer. The audit added one: kernel counters of each container's netns. `inside.py` could record `/proc/net/snmp` before and after the subject in the same netns at no cost. | Suggestion for the kit, not applied |

The results file's "Before any public use" list now also names `audit/witness/audit-sandbox.json`, which
holds the dev laptop's LAN and private-network addresses (they were connect targets).

## Check 1: the sandbox had no route out

**From inside** (`audit/audit_inside.py`, run 09:00:52Z to 09:00:54Z): one container with the probe's exact
flags (`--network none`, uid 1000, `--cap-drop ALL`, `no-new-privileges`, `--read-only`, the probe's
`resolv.conf`), before any listener started. Output: `audit/witness/audit-sandbox.json`.

| Check | Result |
|---|---|
| Interfaces | `lo` only |
| IPv4 routes | 0 |
| IPv6 routes | 3, all on `lo`: `::1/128` plus two reject entries (flags `0x00200200`) |
| Capabilities | `CapEff`, `CapPrm`, `CapBnd` and `CapAmb` all 0; uid 1000 |
| Proxy variables | none |
| `/run`, `/var/run` | `lock` only (no host sockets) |
| `getaddrinfo(collector)` before the fake DNS started | `EAI_AGAIN` (-3) |
| TCP connects: 1.1.1.1:443, 8.8.8.8:53, 20.184.175.9:443, 20.184.175.13:443, docker0 172.17.0.1:80, bridge <the dev laptop's second Docker bridge address>:443, the dev laptop's LAN address :443, the dev laptop's private-network address :443 | all `errno 101 ENETUNREACH` |
| IPv6 TCP connects: [2606:4700:4700::1111]:443, [2001:4860:4860::8888]:53 | both `ENETUNREACH` |
| UDP sends: 1.1.1.1:53, 8.8.8.8:53, [2606:4700:4700::1111]:53 | all `ENETUNREACH` |

**From the host, during the re-runs** (09:01:32Z, 4 live probe containers):
- Each was `NetworkMode none`.
- Each had its own netns (`net:[4026534036]`, `[4026533621]`, `[4026533979]`, `[4026534093]`; the host's
  is `[4026531833]`).
- `/proc/<pid>/net/dev` listed only `lo`, and `/proc/<pid>/net/route` held 0 routes.

**Kernel-counter witness** (`/proc/<pid>/net/snmp` of each container's netns, read from the host every
second): `audit/witness/*.txt`.

| Arm (audit rep) | Counters after the self-test | Counters at the last sample (after the subject exited) | Change during the subject's life |
|---|---|---|---|
| ort130-envoff (audit2), 09:05:51Z to 09:07:54Z | OutRequests 14, OutNoRoutes 3, TcpActiveOpens 1, UdpOutDatagrams 6 | the same | none |
| ort129-envoff (audit3, not pre-registered), 09:10:11Z to 09:12:14Z | OutRequests 14, OutNoRoutes 3, TcpActiveOpens 1, UdpOutDatagrams 6 | the same | none |
| ort130-default (audit2) | the same start | OutRequests 74, OutNoRoutes 3, TcpActiveOpens 6, UdpOutDatagrams 26 | +5 TCP opens (one per attempt), +20 datagrams, 0 no-route |
| ort129-default (audit3) | the same start | the same as ort130-default | the same |

OutNoRoutes stayed at 3 in all 4 sampled arms. The 3 are the self-test's own public connects. So no subject tried a
hard-coded external address, and that holds even for sockets strace cannot see.

## Check 2: the pre-registration predates the first run

| Event | Time (UTC) | Source |
|---|---|---|
| `PREREG-probe.md` created | 08:40:05Z | `stat` birth time |
| The last write to the pre-registration (the freeze stamp) | 08:45:13.5Z | the probe agent's transcript |
| `PREREG-probe.md` mtime | 08:45:16.67Z | `stat` |
| `run-all.sh` preflight | 08:45:25Z | `probe/runs/preflight.txt` |
| First subject launch (posctl-r1 `t_launch`) | 08:45:31.51Z | `probe/runs/posctl-r1/result.json` |

- The current sha256 is `<withheld, see README.md>`. All 18
  `result.json` files record it, and so do the audit's re-runs at 09:01Z and 09:05Z, so the file has not
  changed since.
- The dry run recorded `<withheld, see README.md>`, the text as it stood before the freeze, as expected.
- The lane transcript shows no onnxruntime import before the freeze:
  - uv installs at 08:34Z (an install, no import)
  - a binary grep
  - an 08:35:16Z container that exited with "python: not found"
  - an 08:38Z strace test with plain CPython
  - the 08:44Z dry-run posctl, which never imports onnxruntime
- The lane made no ssh or scp call and never named the inference box (0 matches for its name in its Bash calls).

## Check 3: the logger can see an attempt (positive controls in the same sandbox)

Run by `audit/audit_inside.py` through the probe's own listeners (imported from `probe/inside.py`), each
client under the bundled strace:

| Control | Fake DNS log | TLS / HTTP log | strace |
|---|---|---|---|
| `urllib` HTTPS to `https://mobile.events.data.microsoft.com/OneCollector/1.0/` | A + AAAA for the collector | SNI = the collector, 1558-byte ClientHello, ALPN `http/1.1` | connect 127.0.0.1:53, connect 127.0.0.2:443 |
| `urllib` HTTP to the same name | A + AAAA | Host = the collector, `GET /OneCollector/1.0/ HTTP/1.1` | connect 127.0.0.1:53, connect 127.0.0.2:80 |
| Raw connect to the hard-coded collector IP 20.184.175.9:443 | nothing | nothing (0 listener events) | connect 20.184.175.9:443 = `ENETUNREACH` |

The kit's own posctl, re-run as audit1, logged the lookup and the SNI at 5.024 s, like the original 3/3.

## Check 4: the audit's re-runs

- The wheels were fetched fresh from PyPI on the host, outside the sandbox: the 1.30.0 set at 08:59:36Z and
  the 1.29.0 set at 09:09:56Z.
- All 8 files of the 1.30.0 set and the 1.29.0 onnxruntime wheel have the same sha256 as
  `probe/runs/wheels-SHA256SUMS`, and the pybind `.so` hashes match `probe/runs/ort-pybind-SHA256SUMS`.
- The venvs were built offline in a `--network none` container.
- Every arm except the last pair ran from the archived kit, unmodified.
- The 1.29.0 env-off arm is **audit-only**. It is one added `case` line in a scratch copy of `run-arm.sh`
  (`audit/run-arm-audit-arm.diff`), and it records `prereg_sha256: none`.

| Arm | Audit rep | Exit | Self-test | Lookups | TLS attempts at (s) | inet syscalls | Files | Queue rows | Original (3 reps) |
|---|---|---|---|---|---|---|---|---|---|
| posctl | audit1 | 0 | pass | 2 | 5.0 | 2 | none | – | the same |
| ort130-default | audit1 | 0 | pass | 10 | 9.2, 14.8, 24.5, 45.0, 78.9 | 15 | deviceid, db, .ses, mat-debug-15.log | 12, 4 canaries | the same shape |
| ort130-default | audit2 | 0 | pass | 10 | 9.2, 14.0, 24.3, 42.5, 79.2 | 15 | the same four | 12, 4 canaries | the same shape |
| ort130-envoff | audit1 | 0 | pass | 0 | none | 0 (not even `socket()`) | none | no DB | the same |
| ort130-envoff | audit2 | 0 | pass | 0 | none | 0 | none | no DB | the same |
| ort130-apioff | audit1 | 0 | pass | 10 | 9.2, 12.6, 19.5, 42.3, 76.2 | 15 | the same four | 3, no canary | the same shape |
| ort129-default | audit3 | 0 | pass | 10 | 9.2, 12.2, 19.4, 34.4, 69.1 | 15 | the same four | 12, 4 canaries | the same shape |
| ort129-envoff (not pre-registered) | audit3 | 0 | pass | 0 | none | 0 | none | no DB | not run originally |

Raw files: `audit/runs/<arm>-raudit<n>/`. Re-derived facts: `audit/verify-audit.txt`.

## Check 5: every figure matches the raw logs

`audit/verify.py` re-derived the figures from each run's `strace.log`, read directly rather than from
`result.json`'s parsed copy, and from the listener logs, `marks.json` and an `immutable=1` open of each queue
copy. Output: `audit/verify-orig.txt`.

| Figure in PROBE-RESULTS.md | Raw value (18 runs) | Match |
|---|---|---|
| Lookups and attempts per run (10 / 5; posctl 2 / 1; silent arms 0 / 0) | the same; every TLS accept pairs with a strace `connect` to 127.0.0.2:443 within 1 ms | yes |
| inet syscalls 15 / 2 / 0, never to anything but 127.0.0.1:53 and 127.0.0.2:443 | the same; `other_inet` empty in all 18 | yes |
| Silent arms: 0 network syscalls of any kind | 0 non-exit lines in all 6 strace logs (no `socket()` at all) | yes |
| First attempt 9.2 s; 5 attempts; the per-run times in the table | 9.155 to 9.232 s; every time in the table matches to 0.1 s | yes |
| Gaps 3.4–5.7, 8.1–11.3, 12.5–22.5, 37.2–46.9 s; last attempt 73.8–90.3 s | 3.42–5.74, 8.12–11.26, 12.54–22.53, 37.15–46.91; last 73.78–90.31 | yes |
| deviceid at 0.14–0.22 s; `.ses` at the same moment; the DB mtime 120.22–120.29 s | 0.141–0.221; `.ses` within 1 ms; 120.221–120.292 | yes |
| Queue: 12 rows (the event list as stated) / 3 rows; `retry_count` 0 | the same names in every run; retry 0 in every row | yes |
| The four canaries in `SessionCreation` in A1/A4 (6/6); none in A3 | yes, only in the `SessionCreation` row; none in apioff | yes |
| `c:SHA256(deviceid)`, the `.ses` UUID and the interpreter path in every row; no raw deviceid, machine-id or hostname | true in all 9 telemetry runs | yes |
| ClientHello 469 bytes, no ALPN, TLS 1.3 + 1.2 | the same in all 45 attempts | yes |
| A + AAAA "in the same millisecond", on a separate thread; nscd probe only with the first lookup | spread ≤ 1 ms; DNS pids 64+ (the main pid is 15); AF_UNIX only at about 9.2 s | yes |
| Exit 0.018–0.022 s after the final mark | ORT arms 0.018–0.022; posctl 0.005–0.006 | ORT arms only (N-3) |
| 18/18 self-tests pass, rc 0, 0 events outside the window, 0 listener errors | the same | yes |
| Host facts (24 CPUs, Core Ultra 9 290HX Plus, 64203440 kB, Ubuntu 26.04.1, kernel string, Docker 29.1.3, strace 6.19, image `sha256:88200866…4171`) | re-read at 09:0xZ, the same | yes |
| Packages identical across the venvs except onnxruntime; CPython 3.12.13 | the same in every `result.json` | yes |
| `ProcessInfo` strings (CPU model, Debian osDescription, Desktop, Unmetered/Wired, `EVT-Linux-C++-No-3.10.173.1`, tenant `o:5ad963bd…`) | present in `ort130-default-r1/queue/strings.txt` | yes |
| Only 3 files name the host (`<hostname withheld>`) | the same; no host machine-id and no host deviceid anywhere in `probe/runs/` | yes |

## Check 6: the claims stay inside what was measured

- The verdict lines, the "does NOT show" list and the INFERENCE labels are sound.
- Two claims reached past the data, I-1 and I-2, and both are fixed.
- Three wordings over-generalised, N-2, N-3 and N-4, and all three are fixed.
- The payload claims are correctly limited to the local queue, never the wire.
- The 9 s charging explanation is correctly marked INFERENCE. `ADP1/online` read `1` again at 09:0xZ.

## What this audit did not check

- It did not check delivery, or the bytes on the wire (the probe does not claim either).
- It did not re-run 1.28.0 or 1.29.0 api-off, or any second rep of the audit-only 1.29.0 env-off arm.
- It did not check the recon's source citations (`RECON-onnxruntime-sources.md`), beyond the lines the results
  quote.
- It did not trace what wrote the dev laptop's own queue WAL at 07:54Z (N-9).

## Housekeeping

- The audit scratch (465 MB at peak, on tmpfs) was deleted at 09:14Z.
- `docker ps -a` shows no `probe-` container (0).
- The dev laptop's own onnxruntime queue has the same mtimes as before (`.db` 01:00:47Z, `-wal` 07:54:37Z,
  `-shm` 07:54:52Z).
- The audit's network use was the two PyPI wheel downloads on the host, and nothing else.

**Files:**
- `PROBE-AUDIT.md`: this report.
- `PROBE-RESULTS.md`: the mechanical fixes above, plus a dated audit addendum at the end.
- `audit/`: `audit_inside.py`, `verify.py`, `verify-orig.txt`, `verify-audit.txt`, `run-arm-audit-arm.diff`,
  `wheels-SHA256SUMS`, `ort-pybind-SHA256SUMS`, `runs/` (8 audit runs), and `witness/` (the sandbox JSON
  and the netns counter logs). 660 KB in total.
