# PRE-REGISTRATION: the no-egress probe for onnxruntime's phone-home

Written 2026-09-28, 08:39Z to 08:45Z (UTC, `date -u`), on the dev laptop, **before any onnxruntime process was started
by this probe**. Frozen at 08:45Z. Its sha256 is recorded in every run's `result.json` (`prereg_sha256`), so any edit after the
first run is visible. Amendments after freeze go at the bottom, dated, and never rewrite the text above them.

## Question

Does onnxruntime (>= 1.29, default settings) try to reach Microsoft? When, and how often? Does
`ORT_DISABLE_TELEMETRY=1` stop every attempt? Nothing may leave the box while we find out.

## Instrument

Each run happens inside a Docker container started with `--network none`. That container has one network
interface, `lo`, and no routes, so nothing can leave the box. Inside it:

- **A logging DNS server** on `127.0.0.1:53` (UDP and TCP, Python stdlib). It logs every query (time, name,
  type). It answers every A query with `127.0.0.2`, and every other type with NOERROR and no answer. The
  container's `/etc/resolv.conf` is replaced with `nameserver 127.0.0.1`.
- **A logging TLS listener** on `127.0.0.2:443`. It logs each connection and parses the TLS ClientHello's
  SNI and ALPN, then closes **without** completing TLS. A plain listener on `127.0.0.2:80` logs any HTTP
  request line and Host header.
- **strace** (host binary, bundled into the container) wraps the subject:
  `-f --seccomp-bpf -e trace=%network`. It records every `connect()`/`sendto()`/`sendmsg()` of the whole
  process tree, to any address. That includes a hard-coded IP that the DNS log would never see (such a
  connect fails with ENETUNREACH, but strace still logs it).
- **A self-test runs before every subject.** Any failure voids the run:
  - TCP connects to `1.1.1.1:443`, `20.184.175.9:443` and `20.184.175.13:443` must fail with ENETUNREACH.
  - `getaddrinfo("probe-selftest.example")` must return `127.0.0.2`, and the DNS log must show the query.
  - A TLS connect with SNI `probe-selftest.example` must appear in the TLS log.
- **Once before the probe, from the host:** `curl` from a `--network none` container (the local
  `curlimages/curl` image) to `https://1.1.1.1/` must fail. The output is recorded in `runs/preflight.txt`.

Container hardening: `--user <uid>:<gid>` (not root), `--cap-drop ALL`, `--security-opt
no-new-privileges`, `--read-only` root, private tmpfs `/tmp`, a throwaway `HOME=/tmp/home`, and a fresh
random `/etc/machine-id` per run (the host's is never shown). Image `debian:12-slim`
(`sha256:88200866…`). CPython 3.12.13 is mounted read-only at `/probe/python`. Docker gives the container's
netns `net.ipv4.ip_unprivileged_port_start=0`, so the unprivileged listeners can bind 53 and 443 (verified
08:36Z: the value read `0` inside such a container).

### Amendment 0 (before any run): why Docker, not `unshare`

The lead's design asked for `unshare --user --map-root-user --net --mount`. On the dev laptop that fails:

- `kernel.apparmor_restrict_unprivileged_userns = 1`, and
  `unshare --user --map-root-user --net --mount …` → `unshare: write failed /proc/self/uid_map: Operation not
  permitted` (08:30Z).
- `bwrap --unshare-all` works, and brings `lo` up. But its AppArmor profile strips capabilities from its
  children (`/etc/apparmor.d/bwrap-userns-restrict`: `audit deny capability`).
  - Inside it, `ip_unprivileged_port_start` reads `1024`.
  - Binding 53 and 443 gives `[Errno 13] Permission denied` (8443 binds fine).
  - Writing the sysctl gives `Permission denied` (08:32Z).
- I did **not** use `nsenter` into bwrap's user namespace to regain capabilities. That would sidestep the
  AppArmor userns restriction on purpose, and that is the operator's call, not this lane's.
- An operator's account is in the `docker` group, and Docker's `--network none` gives the same isolation: only `lo`, no
  routes. Its netns default of `ip_unprivileged_port_start=0` lets the listeners bind 53 and 443 as uid 1000
  with every capability dropped.

The container is not an unprivileged user namespace: the Docker daemon runs as root. Isolation comes from
the kernel netns, the same mechanism, and strace is an added witness the lead's design did not have.

## Subject

Wheels are fetched from PyPI **outside** the probe, with the network. Venvs are then built from those wheels
**inside a `--network none` container** (`pip --no-index`), so even the install makes no network call. The
same pins apply to all three: `onnx==1.23.0`, `numpy==2.5.3`, plus whatever those resolve to (recorded per
venv in the results).

| Venv | onnxruntime |
|---|---|
| `v1300` | 1.30.0 |
| `v1290` | 1.29.0 |
| `v1280` | 1.28.0 (control: no POSIX telemetry, per recon) |

**Workload** (`probe/subject.py`), in this order:

1. `import onnxruntime`
2. **(A3 only)** `onnxruntime.disable_telemetry_events()` immediately after the import
3. build a one-node `Add` model in-process with `onnx.helper`. It is loaded from bytes, so there is no file
   name. It carries three canaries:
   - graph name `probe_graph_canary`
   - producer `probe-kit`
   - metadata `probe_meta_key` = `probe_meta_value_canary`
4. create one `InferenceSession` on `CPUExecutionProvider`
5. call `run()` once
6. sleep 120 s, then exit 0

The subject's environment is built from scratch:

- It contains only `PATH`, `HOME=/tmp/home`, `LANG=C.UTF-8`, `PYTHONDONTWRITEBYTECODE=1`, and the arm's
  variable.
- So no `CI`-style variable can suppress telemetry by accident, and no proxy variable is set.

## Arms (3 reps each; reps run one after another, the 6 arms of a rep in parallel, each in its own container)

| Arm | Venv | Switch |
|---|---|---|
| A1 `ort130-default` | v1300 | none |
| A2 `ort130-envoff` | v1300 | `ORT_DISABLE_TELEMETRY=1` in the subject's environment |
| A3 `ort130-apioff` | v1300 | `disable_telemetry_events()` right after import; no env |
| A4 `ort129-default` | v1290 | none |
| A5 `ort128-default` | v1280 | none |
| P0 `posctl` (instrument positive control) | v1300 | does **not** import onnxruntime. Python itself resolves `mobile.events.data.microsoft.com` and opens TLS to it on port 443 at about 5 s. It then sleeps to 30 s |

## Measures, per run

- **lookup**: a DNS query logged by the fake server, with its name and type.
- **collector attempt**: a TCP connection accepted on `127.0.0.2:443`, with its SNI. Cross-checked against
  strace `connect()` to `127.0.0.2:443`.
- **any network attempt**: any strace `connect`/`sendto`/`sendmsg`/`sendmmsg` to an AF_INET or AF_INET6
  address. AF_UNIX (for example glibc's nscd probe) is local and is listed separately.
- All times are relative to `t_launch`, the harness's wall clock just before it starts the subject.
- **Files**: the full listing of `/tmp` (which holds `HOME`) before and after, and the queue DB rows after
  exit (count, latency, retry_count, event names, canaries).

## Predictions (primary: these decide the verdict)

- **P1, A1 `ort130-default`**: at least 1 lookup of `mobile.events.data.microsoft.com` **and** at least 1
  collector attempt with SNI `mobile.events.data.microsoft.com` within **N = 60 s** of `t_launch`, in 3/3 reps.
- **P2, A2 `ort130-envoff`**: **0** lookups, **0** collector attempts and **0** non-AF_UNIX network
  syscalls from `t_launch` to exit + 3 s, in 3/3 reps.
- **P3, A3 `ort130-apioff`**: the same as P1. The runtime API leaves the uploader live, and ProcessInfo is
  already queued.
- **P4, A4 `ort129-default`**: the same as P1.
- **P5, A5 `ort128-default`**: the same as P2.
- **P0, instrument**: in every rep, `posctl` shows at least 1 lookup and at least 1 SNI
  `mobile.events.data.microsoft.com`. Every run's self-test passes.

A prediction holds only if it holds in every rep. One contrary rep refutes it, and the refutation is reported
as such.

## Secondary (directional, reported but not deciding)

- **S1**: in the P1, P3 and P4 arms, the first collector attempt comes at **5 to 40 s**. Recon's BEST_EFFORT
  timer is 18 s, or 9 s when the SDK reads the power source as charging.
- **S2**: the P1, P3 and P4 arms make **more than one** collector attempt in 120 s, because failed uploads are
  retried. No count is predicted.
- **S3, files**:
  - A1, A3 and A4 create `$HOME/.cache/Microsoft/DeveloperTools/.onnxruntime/{deviceid,onnxruntime.db}`
    and `/tmp/.ses`.
  - A2 and A5 create none of them.
- **S4, queue after exit** (the uploads failed, so the events should persist):
  - A1 and A4 hold session-level events, and at least one payload carries `probe_graph_canary`.
  - A3 holds `ProcessInfo` but **no** `SessionCreation` and no canary.
- **S5**: no payload contains the raw `deviceid` UUID, the run's raw machine-id, or the container hostname
  `ortprobe-host`.

## Instrument dry run before the freeze (no onnxruntime involved)

This ran on 2026-09-28 at 08:44Z. Output is in `probe/runs/dryrun/`.

- `preflight.sh` PASS:
  - curl to `1.1.1.1` and to `20.184.175.9`: exit 7, "Could not connect to server".
  - curl to the collector name: exit 28, "Resolving timed out".
- `run-arm.sh posctl 0`:
  - The self-test passed.
  - The fake DNS logged `mobile.events.data.microsoft.com` A and AAAA at 5.025 s.
  - The TLS listener logged SNI `mobile.events.data.microsoft.com` at 5.026 s (1539-byte ClientHello).
  - strace logged `connect` to `127.0.0.1:53` and to `127.0.0.2:443`, plus glibc's two AF_UNIX nscd probes (ENOENT).
  - The subject exited 0 at 30.03 s.

## Void and flag rules

- **Self-test failure:** the run is void. It is kept in `runs/` and marked void, then re-run once with a note.
- **Subject exit code other than 0:** the run is flagged but still counted. Attempts are attempts.
- **Harness or container error before the subject starts:** void.

## What this probe cannot show (declared in advance)

- It shows **attempts**: DNS queries, TCP connects and ClientHellos. It cannot show what the real collector
  would have accepted, or what a real TLS session would have carried.
- The payload is read from the **local queue DB** after exit, not from the wire.
- The container is not a host:
  - `/etc/os-release` is Debian 12, and the machine-id is fake.
  - `/proc/cpuinfo` and `/proc/meminfo` are the host's.
  - Behaviour that depends on host daemons, `/run`, or real network-type detection may differ.
- 120 s per run. Longer-horizon behaviour (retry drop after 5 failures, 10-min RuntimePerf cadence) is out
  of scope.
