Research, with receipts.
We build software that runs on machines you own — a game, a rules companion, a walking-tour guide, a private document AI for organizations — and every model, sampler, and pipeline choice inside it is decided by a bench, not a vibe. This hub collects the research those decisions produced — real experiments on real product content, receipts under every claim. When a bet loses, the loss is published too. Beside the benches sit field guides to how our labs work, galleries of what the painters made, and workshop notes — each with its kit.
New here? In plain words: we test AI models — every one we have tried is on the model roster, and every card and machine we have measured is on the hardware roster — before our products are allowed to use them, and publish every result with the data behind it — including the failures. Each article opens with a short glossary of the few terms it leans on. A good first read is the newest piece below, or how we work — who writes this, and how.
The exhibits
The exhibits are the workshop’s bench record — model auditions, serving benches, and operational memoirs, each carrying its receipts — thirty-four of the sixty-three with a full data kit, and the index names which. The rows say what each one settled — newest first.
2026-09-22 A laptop, asked the desktop's questions One RTX 5090 Laptop part, 24 GB, built into an off-the-shelf gaming laptop, put through the same language and render bench this shelf ran on a 3090, a 3090 Ti and a pair of 3080s. It was measured three times: held to its 95-watt default; with the maker's Dynamic Boost left to float the budget, which it held between 140 and 150 W; and the way the machine boots, boost and a clock lock together, which changed nothing under load and four watts at idle. The finding is a ratio, not a deficit: a part allowed 95 watts, where the desktop cards beside it are allowed 350 to 450, does most of the work on the model this workshop runs, loses a third of it on a dense one, and gets the third back the moment it is allowed 150. Every watt on the page is the card's own draw and not the wall's, the fan was held at maximum by hand, and nine things the page cannot say are written down. The bench 2026-09-21 Four homes for one reranker: what RuleSage's slowest stage cost at each stop On the morning of 2026-09-17 an operator at a game table waited nearly twenty seconds for the first word of a rules answer and thought the app was broken. It was not. One small stage of the answer, a cross-encoder that re-reads the shortlist before anything is written, had moved to a rented server whose processors lack the instruction set its weights were built for. This page follows that stage across four machines, three of them in five days: two rented VPSes, a mini PC, and a graphics card that was already in the house. Every millisecond we measured at each stop is here, with the gate we ran before the last move, the first word a player waits for after it, the traps that bit on the way, and what is still unmeasured. The bench 2026-09-21 A 3090, in its own words Fiction, and the byline says who wrote it: a used EVGA GeForce RTX 3090 XC3 Ultra, 24 GB, tells its three lives in the first person: a year of mining, a sky across two monitors, and now the minds it serves.
2026-09-20 The indifferent prediction On the night OpenAI published a resolution of the Navier–Stokes existence and smoothness problem, an operator — one person, the same one throughout this page — and a model looked at the one figure in the post and each said what they saw. Nothing on this page is this workshop's measurement: it is two readings of somebody else's result, the story the picture arrived into, and the thought the night ended on — about a third kind of mind, one that would learn the universe without ever learning us. Workshop notes 2026-09-18updated
2026-09-22 RTX 3090 vs RTX 3090 Ti: an inference and diffusion bench One EVGA GeForce RTX 3090 XC3 Ultra 24 GB and one EVGA GeForce RTX 3090 Ti FTW3 Ultra 24 GB, in the same x16 slot of the same desktop, on the same supply and the same script, 17–18 September 2026. Numbers only: what each writes, what each holds, what a picture costs, how hot each ran, and what a power cap takes away. The bench 2026-09-17updated
2026-09-21 Two 3080s against one 3090, and the cap decides Two GeForce RTX 3080 10 GB cards bought in 2021, in a desktop of the same vintage, against one GeForce RTX 3090 24 GB in the same box, on the same supply, through the same script. The single card is faster on three of the four things measured and uses about half the energy; the pair reads a long document in two thirds of the time and costs about two fifths as much. The bench 2026-09-16 A Short History of Mistral A French lab that put its first model on a torrent, every open-weight release since with its licence, and what its 24B model does on a 3090 and a 96 GB workstation card. Workshop notes 2026-09-16 Ask About This Page Every article on this hub has a room now, where a model on our own machines answers from that page’s own words and leaves its first draft on the page, marked where the checks took something out. This is how it is built, what it keeps, what it refuses, and what it got wrong on the way here. Workshop notes 2026-09-15updated
2026-09-21 A Dinner Party for the Dead A dinner party for three people who are dead and do not know it, hosted by you. Six guests are on the shelf, played by two models on our own machines.
2026-09-18 The cost to purchase and run a very capable home AI rig A consumer graphics card in a consumer desktop, nothing server-grade anywhere in the box: the rig costs $3,198.86 with one card and $5,048.85 with two on 13 September 2026, before tax. Running it costs about $18 a month with one card and $31 with two if every day puts them to work for six hours; with the cards bought used at August's average completed sale, the rig is $2,443.31 with one and $3,537.75 with two.
2026-09-12 How the print lab works A darkroom that develops three pictures a day for your household, a board that fills up and then finishes for good, and an archive where nobody has a name.
2026-09-11 Six Worlds for the Print Lab Six worlds, painted large, for a print lab that is now open — a darkroom out back, a wall out front.
2026-09-18 The Beat Lab, asked and answered Every question the Beat Lab’s genie can answer, answered here too — the same twenty-eight answers, out of the same file, read by a page instead of a genie, each one set beside the sentence of the field guide it was written from. The field guide 2026-09-03 Ten Minutes with Living Artists A genre instead of a composer: 159 tracks of Creative Commons minimal techno, the same ten-minute recipe that taught three dead composers, and the one ear in the house preferred the model with nothing added — both times. The field guide The bench 2026-09-03 Listen for Yourself Every clip of the fourth adapter’s four rounds — sixty-five renders, raw and brought to a common loudness — so you can hear what one listener heard and decide for yourself. The field guide 2026-09-02 Half an Hour with Dead Composers Two more dead composers taught the same way — a hundred and six marches, then Bach — and then blending composers with linear algebra, where the obvious method is wrong, the right one costs a measurable one percent, and a sentence already staged for publication had to be taken back before you ever saw it. The field guide The bench 2026-09-01 Teaching a Music Model Chopin in Five Minutes Nineteen public-domain Chopin recordings, four minutes and forty-nine seconds of training on a five-year-old graphics card, and a dial you can hear — five clips on the page that differ from each other in exactly one way. The field guide The bench 2026-08-31 The Same Sixteen We put glm-5.3-flash — ollama’s new 320B-total cloud model — on two frozen house instruments; it tied our local 12B and 27B at 16/19, and the finding is about the ruler. The bench 2026-08-31 Two Hours at 12 tok/s — On Battery Take the graphics card away and the 30-billion-parameter model answers faster than the 9-billion one — on all three machines. The bench 2026-08-29 The Instrument Travels Carry a frozen exam from a 96 GB workstation card to a 24 GB consumer one and the model it ranked two weeks ago ranks again — at the floor, on a different build, with nothing new seated. The bench 2026-08-28updated
2026-09-04 Nine Worlds, One Dog The founder's dog takes the same nap under the same tree in all nine painted worlds — same dog, same curl; only the world changes.
2026-08-28 What 150 Watts Buys Turn a 600-watt card down by a quarter and one production model pays two percent of its speed; the other never notices. The bench 2026-08-27updated
2026-09-15 How the Beat Lab works A drum machine that lives in your browser, a genie that runs on our own hardware, and a board where nobody has a name. The field guide 2026-08-27 The Typist and the Developer Two ways a machine writes: one word at a time, or every pixel at once — sixteen times over. The field guide The bench 2026-08-26 A Rig Your Friend Already Owns We re-rendered our game's actual art — same prompts, same seeds, same checkpoint files — on a several-year-old consumer RTX 3090 and a gaming laptop, paired render-for-render against our 96 GB card's own archived output. Two of the three runs happened in the same physical computer. This page is the stopwatch, told honestly — including the one number we don't yet trust.






The bench
2026-08-25updated2026-09-04 The Dog, the Dice, and the Painter: How a Small World Draws Itself A photograph of a real dog walks into a painted fishing cove and comes out as a painted dog the world will remember. This is the whole pipeline between those two facts — the file, the dice, and the painter — one station at a time, with no scores and no instrumentation.



The field guide
2026-08-24 The new-arrivals diffusion bench One small fishing village, painted thirteen ways — Sorrowmoor Cove (and some other worlds built on RealKeep) sits for thirteen open-weight painters: 3,588 published renders across seventy-five galleries, every hero hand-picked, every licence read first-hand, and the two painters whose licences refuse counted in the open rather than quietly dropped.



2026-08-21 Everyone on the payroll, three at the table — dense, MoE, and active weights Model names grew a second number — 30b-a3b — and it is the one your card and wallet care about. The payroll and the meeting, in plain English, measured on our own machines: a 25.8B mixture out-writing its 11.9B dense sibling while holding twice the memory — and the robber asks both architectures. The field guide 2026-08-19updated
2026-08-21 Reading is fast, writing is slow — why AI answers arrive word by word Paste a whole essay into a chatbot and it swallows the thing at once — then answers one word at a time. That asymmetry is physics, not theater: reading and writing in plain english, measured on our own card — and the measurement argued back, which turned out to be the best part. The field guide 2026-08-18updated
2026-08-21 Three librarians and a careful reader — how RuleSage finds the right page Before RuleSage answers, something has to find the right page in a book the size of a small novel — three searchers, one referee, one careful reader, and none of them is the model that writes your answer. The whole journey, in plain english, with the shipped numbers. The field guide Workshop notes 2026-08-17updated
2026-08-20 qwen3.8:27b across four house benches: one seat filled, one floor missed qwen3.8:27b landed on our box the day after it shipped, and we sat it in the exams we already had rather than designing anything for it — it filled a screening seat that had never been filled, and it missed a floor in the judge seat. The bench 2026-08-17updated
2026-09-08 The map nobody picks up — and whether it helps when it arrives. Two camps have argued about a map file for machines for two years, mostly without receipts — no AI crawler asked us for ours in thirty days of logs, so we measured whether it helps when it arrives anyway. The bench Workshop notes 2026-08-15updated
2026-08-20 The open call — a kid, an elder, and a tired parent walk into the cove. Twenty language models were handed the same frozen moment of a small fishing town and asked to speak as its people, then read blind by seven judges from six rival families. The bench Workshop notes 2026-08-15updated
2026-09-13 Where your question goes — a plain-english walk. Our two published privacy promises, unpacked for people — not programmers — by following one question door to door. Workshop notes 2026-08-14updated
2026-09-10 How we work — who writes this, and how. The honest answer to the fair question: the workflow that makes every exhibit — humans and agents, named plainly, with the receipts. Workshop notes 2026-08-14updated
2026-08-20 The move — the same yardstick, before and after. Our apps left the gaming laptop for a real server — the same frozen bench measured both eras, honestly. The bench Workshop notes 2026-08-13updated
2026-08-20 The chair trials — five fresh exams for the thirty-billion class. Nine models, eleven arms, five all-new exams in one day — on a box that never stopped serving its real users. The bench 2026-08-13 The narrator’s chair, refused. The August class’s best voice sat the cove’s actual chair — and was refused at the gate. The bench 2026-08-13updated
2026-08-14 The outside judges — four rival labs re-check our work. Four rival frontier labs re-judged every sealed round we had published — blind, byte-identical. The bench 2026-08-12updated
2026-08-20 The August arrivals — the fresh class sits the house exams. Four open-weight models landed in one week, so we gave them the two exams we already had — one frozen since July, one rebuilt in its shape. The bench 2026-08-11 The diffusion research bench. What a fully local image pipeline delivers for a living world — head-to-heads, ladders, cold starts.



The bench
2026-08-11updated2026-08-19 The seat trials — how a local model earns a chair. Before a model earns a seat inside our products, it sits this trial. The bench 2026-08-11updated
2026-08-13 The voice trials — how the narrator earned its voice. Twenty-four model tags across three hardware eras, hunting a voice that can inhabit a character. The bench 2026-08-11updated
2026-09-04 The licence ledger — every model, read first-hand. Every model we run or bench, its licence read from the actual text — the page the others cite instead of restating. Workshop notes
The roster
Looking for one model rather than one bench? The model roster lists every model that has appeared in a published exhibit here — Muse Glimmer, Qwen 3.6 and 3.8, Gemma 4, Nemotron 3.5 Lightning, OLMo 3.1, Granite 4.1, Kimi K3, DeepSeek V4, Mistral Large 3, GPT-5.5, Claude Opus 5, GPT-6 Astra, Claude Fable 5.1, the FLUX.1 painters and the rest — with the exhibits each one appears in, that round’s own verdict quoted word for word, and its row in the licence ledger. It is an index, not a bench: it is regenerated from the exhibits’ data kits and settles nothing on its own.
The lab protocol
Every exhibit on this hub was produced under the same discipline. None of it is a standard anyone handed us — we wrote each law down after a bench told us something we didn't want to hear, and we've kept to them since:
Quality gate first — the pass bar is written down before the experiment runs.
Measured verdicts — a claim without a receipt is not published; every figure is a
count of rows in a named artifact, and raw rows go out on request.
Self-refutation honored — when the data kills our favorite, the favorite dies.
Real content — benches run on the products' own material, never toy prompts.
Losses published — a rejected bet gets its numbers shown, not a memory hole.
Nothing leaves the box unnamed — benches run on our own authored content, on our
own hardware. When a bench reaches a hosted endpoint (the seat trials' reference ceiling
does), the page says so plainly, and no customer or user data rides in any prompt.
Hardware by class — machines are named by VRAM class, not model; within-rig comparisons
are exact, cross-rig extrapolation is approximate. Specific specs are available on request —
drop a line.