The notes — what they are and how they work
Uncensored models
exhibit seventy-one The notes
Published 2026-10-05 (UTC)
A small (human) team and a fleet of AI agents.
An "uncensored" AI model is one that somebody has altered so that it refuses far fewer requests than the version its maker released. The change is made to the model's files, or to the way it is run. This page explains where a model's refusals come from and the three ways people remove them. It then takes one popular uncensored copy of a large model, shared online, and checks, file by file, what actually changed. It ends on a practical question: what to keep beside such a model on a computer with no internet connection. No code, no settings, no steps.
One of two pages on "uncensored" models: this page · the companion bench test (designed, not yet run)
the short version
An "uncensored" model is an AI model, or a way of running one, changed so that it declines far less often than the version its maker shipped. In a model you download and run yourself, refusing is a habit taught after the main training. Research since 2024 finds that habit carried largely by one pattern in the model's working numbers, which researchers call a direction (later work finds more than one). People take it out in three ways: they train a model without the habit, or over it; they make one permanent edit to the model's learned numbers (its weights) that removes that direction; or they act on the model while it runs, which stops when you switch it off. None of the three makes a model smarter or gives it secret knowledge. We read one popular "uncensored" copy of Qwen3.8-Flash-Next, a large model that Qwen publishes for anyone to download. Both sit on Hugging Face, the site where people share AI models; the uncensored copy's page there was created two days after Qwen's own. We read it by the fingerprints (hashes) the site publishes for every large file, without downloading a byte of its weights. In the uploader's full-precision copy, 80 of the 131 files are byte for byte Qwen's own, and all 51 that changed hold a part the uploader's model card, the description page beside the files, says it edited. So as far as public evidence reaches, the copy is what its card says it is. The serious refusals go too: by the uploader's own figures, the edited model declines harmful requests far less often than the original. We also found a program for running models on your own computer whose optional "speed" setting is, byte for byte, a refusal-removal file published elsewhere; its setup question says what it does, and its name does not. And if a model is your only source of knowledge on a machine with no network, we say what we would build: the model beside offline references it can cite.
8,306 words, about 38 minutes to read.
The summary is this page’s own.
If you're new here: strata→signal is a small workshop (plus a friendly dog with a white patch) that builds on its own machines and writes up what it measures. This page runs no model of our own, on purpose. Every figure on it was read from the source named beside it, on 2026-10-02 (UTC) unless the sources list or the line itself says otherwise, or counted or worked out from those records on the page itself. Most are what Hugging Face, the site that hosts the model files this page reads, publishes about each repository, the online folder that holds one model's files; the rest come from the papers, the licence texts and the uploader's own cards. Each card is cited at a pinned revision: one saved version of a repository, named by a code like e096800036ec, cited exactly because a repository can change. Where we found no source, the claim is not here.
Short on time? The short version above is the page in brief. When the model is the only book on the shelf is the part closest to everyday life, and What to take with you is the page in a few lines; the rest explains how it works and shows the evidence, for anyone who wants it.
Eight words this page leans on. A model's weights (also called its parameters) are its learned numbers, often billions of them (Qwen describes the model this page reads as 125 billion parameters), and they make up nearly all of the bytes you download. A shard is one of the files a large model is split into; the model on this page comes as 131 of them. BF16 is a 16-bit number format; the BF16 copy is the uncompressed one its maker published, which this page calls the full-precision copy. A GGUF is a file in the format read by llama.cpp (an open-source engine for running models on your own computer), usually a compressed copy of that; the compressing is called quantization, and the codes on the files (Q4_K_M, IQ3_M) say how hard it was squeezed; a compressed copy is often called a quant. A file's hash (on this page, its sha256, which Hugging Face publishes for every large file) is a fingerprint computed from every byte in it: change one byte and the fingerprint changes completely, so two files with the same hash are, for every practical purpose, the same file. The residual stream is the model's running state: one working vector per word-piece of the text, which every layer reads from and writes into as a reply is produced. We say the running state below. Each working vector is a long list of numbers, which you can picture as a point in a space with one axis per number; a direction is a way of moving through that space, and the drawing D2 below shows one in two dimensions.
Where the no comes from
If you have ever asked an AI chatbot a plain question and been told "I can't help with that", you have met a refusal, which this page calls the model's "no". A base model, the thing that comes out of the main training run over an enormous pile of text, has no reliable habit of declining; it continues whatever you hand it. The "no" is put in afterwards, in the stage called post-training, and it is put in deliberately; the field's word for this shaping of a model's behaviour is alignment. Instruction tuning with human feedback (Ouyang et al. 2022, the InstructGPT paper) fine-tunes the model (trains it further, on a much smaller set of examples) on "labeler demonstrations of the desired model behavior" and then, by reinforcement learning, on human rankings of the model's outputs. Constitutional AI (Bai et al. 2022) trains "a harmless but non-evasive AI assistant" against a written list of principles. Direct preference optimisation (Rafailov et al. 2023) reaches the same place more simply, and Meta's Llama 2 paper is the worked public example: its Section 4, "Safety", covers safety fine-tuning and red-teaming, attacking a model on purpose to find its weak spots (Touvron et al. 2023). A 2026 thinking model, one that writes out a stretch of reasoning before its answer (Flash-Next is one, and its thinking can be switched off), can carry a further kind: deliberative alignment (Guan et al. 2024) "directly teaches the model safety specifications and trains it to explicitly recall and accurately reason over the specifications before answering".
Sometimes the no arrives by accident. In 2023 the instruction data for many freely downloadable models was generated by asking a hosted chatbot, and that chatbot's refusals came along with its answers. Eric Hartford, whose May 2023 article "Uncensored Models" on his own website is one of the early public accounts of uncensoring, put it in one line: "when the dataset contains answers where the AI is being coy or outright refusing (called Refusals) then the bot learns how to refuse, and under what circumstances to refuse, and how to word the refusals." That accidental no is the very thing his 2023 filtering removed.
There is a third kind of no, which this page is not about: a hosted chat service can put a separate filter model in front of the one you talk to, reading your prompt and its reply. Llama Guard (Inan et al. 2023), "an LLM-based input-output safeguard model", is the published example (an LLM, a large language model, is the kind of model this whole page is about). A model you download and run yourself has no such thing bolted on, which is why its no has to be somewhere inside the files.
Teaching a model to refuse has a known cost: over-refusal, the plain question declined because it rhymed with a dangerous one. Two of the test sets built to catch it: XSTest (Röttger et al. 2023), "250 safe prompts across ten prompt types that well-calibrated models should not refuse to comply with", and OR-Bench (Cui et al. 2024), "80,000 over-refusal prompts across 10 common rejection categories".
Why people want them
The makers' own argument is older than most of the techniques below. Hartford's article has a section headed "Why should uncensored models exist?", and it gives four reasons. Two in his words: "Every demographic and interest group deserves their model", and composability, "To architect a composable alignment, one must start with an unaligned instruct model." Two in our summary of his: that alignment gets in the way of fiction whose characters do evil things, of role-play, and of research and curiosity; and that you should control the software running on your own computer. One of his headings argues that uncensored models should be built, published, maintained and openly available, "for science and freedom and composability and sexy stories and the lulz" (the last word is internet slang for laughs).
On the model this page reads, Qwen3.8-Flash-Next, the uploader of its "uncensored" copy reports over-refusal on XSTest's 250 harmless prompts: with thinking off, 9.6 % refused before the edit and 1.2 % after (by our arithmetic, 24 prompts and 3); with thinking on, the model's default, 0.4 %, one prompt of the 250, both before and after (the uploader's figures, unreproduced, by its own check of each reply's opening words, on the parent card, the card of the uploader's full-precision copy, orcarouter/Qwen3.8-Flash-Next-Uncensored at revision e096800036ec). Who they are made for, by the makers' own tags and words: red teams (the upload is tagged ai-red-team, and the uploader sells a security tier), fiction writers and people who want full control of a model on their own machine (Hartford's reasons), and researchers who study refusal itself. Who actually downloads them, no counter says. One more case gets its own section below: a model kept for when there is no network at all.
The counterweight, in the uploader's own numbers: the same edit takes refusal of harmful prompts "from 64-100% (base) to ~0-3.3%" (the GGUF card, the card of the upload itself, orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF at revision e43d00f4e2b8); the parent card says its refusals are judged "by a rule-based opening-phrase classifier", "indicative, not an LLM-judge / publication-grade number". On this upload, by its maker's own numbers, the edit took out the serious declines and, with thinking off, the silly ones too: one habit, both kinds. One 2024 paper reports that it need not go that way: Wang et al. (2024) remove "a false refusal vector" and report that this "reduces false refusal rate while preserving the model's safety and general capabilities". This upload's edit is not that selective kind; the harmful-prompt figure that opens this paragraph shows as much.
And the licence. Uncensoring changes the weights, not the licence: whatever the licence says still binds the edited copy. Some model licences carry a use policy (Meta's "Llama 3 Acceptable Use Policy" and Google's "Gemma Prohibited Use Policy" are two we read on 2026-10-02); Qwen's carries none. The Qwen Community License 1.0 has two conditions, one about carrying the notice and one about commercial use by businesses that host models or sell AI coding and office assistants, and beyond them it asks only that use comply with applicable law and not infringe anyone's intellectual property. For this model, the no that remains is the law and the person running it.
Three places to touch it
Picture a track in a studio. A model's no is something in the signal, and there are three places you can touch a track: you can re-record it, you can re-cut the master (the finished recording every copy is made from), or you can put a mute on the channel while it plays. Everything below is one of those three.
(a) Re-record it. Two flavours. Never teach it: take a base model and instruction-tune it on a dataset from which the refusals (and, in the 2023 version, answers Hartford judged biased) have been filtered out, so the base model never learns the habit. Hartford's 2023 article is an early account. His "strategy for uncensoring a model is pretty simple. Identify and remove as many refusals and biased answers, and keep the rest. And then train the model with the filtered dataset in exactly the same way that the original model was trained", and he credits "work already done to uncensor Vicuna".
Train it away: take an already-aligned chat model and fine-tune it on a small set of examples that comply. Qi et al. (2023) undid a hosted model's safety training by fine-tuning "on only 10 such examples at a cost of less than $0.20"; Lermen et al. (2023) did it to Llama 2-Chat at 70 billion parameters "with a budget of less than $200 and using only one GPU" (one graphics card), reaching "refusal rates of about 1%"; Yang et al. (2023) name the same thing shadow alignment.
Why does so little work? The papers' answer is that the safety habit is shallow and small: Qi et al. (2024) argue, with evidence, that alignment "adapts a model's generative distribution primarily over only its very first few output tokens" (tokens are the word-pieces of the glossary), and Wei et al. (2024) find the safety-critical regions of the weights are sparse: "about 3% at the parameter level", that is, of the model's individual numbers. LoRA (Hu et al. 2021), the trick Lermen et al. used, which trains a small add-on rather than the whole model, is why such a run can be cheap, not why it works. The cost: the first flavour is a full instruction-tuning run, so what you get is a different model; the second is a small training run, with some drift in everything else.
(b) Re-cut the master: find the direction and remove it from the weights. Arditi et al. (2024) studied 13 open chat models up to 72 billion parameters and found that "refusal is mediated by a one-dimensional subspace": for each model there is "a single direction such that erasing this direction from the model's residual stream activations prevents it from refusing harmful instructions, while adding this direction elicits refusal on even harmless instructions". Once that direction is known, a one-time edit to every matrix (a grid of the model's learned numbers) that writes into the running state removes their ability to express it. No training; permanent; the files on disk are different files. Among people who make and share these edited copies, the usual word for the weight edit is "abliteration"; the paper's own word for its method is "a novel white-box jailbreak method that surgically disables refusal with minimal effect on other capabilities", and its authors read their finding as a warning: "Our findings underscore the brittleness of current safety fine-tuning methods."
This is a simplified picture. Later work refines it: Wollschläger et al. (2025) "uncover multiple independent directions and even multi-dimensional concept cones that mediate refusal", and Marshall et al. (2024) argue, in their title, that "Refusal in LLMs is an Affine Function". The shard census further down this page (D3) shows one such edit's footprint in the actual files.
(c) The mute on the channel: leave the weights alone and act on the running state while the model works. Two flavours again. Add a direction: activation addition (Turner et al. 2023) computes a steering vector from a pair of prompts and works "by tactically adding in e.g. the 'Love' - 'Hate' steering vector during the forward pass"; representation engineering (Zou et al. 2023) is the wider programme; the control vectors that mainline llama.cpp (the project's own, unmodified code) supports are this kind. Remove a direction: Arditi et al.'s own run-time intervention, the same kind of removal as (b), done while the model runs instead of baked into the weights, and gone again when the switch is off.
What this page leaves out: merges with a model that is already uncensored, and prompt-level tricks, which change no weights at all.
And the research pushing back. Circuit breakers (Zou et al. 2024) interrupt models "as they respond with harmful outputs" by acting on the representations rather than on refusal training. Tamper-resistant safeguards (Tamirisa et al. 2024) aim at openly published models whose safeguards "adversaries cannot remove" "even after hundreds of steps of fine-tuning". And a 2025 defence fine-tunes a model to explain itself before declining, spreading the refusal "across multiple token positions"; on three small models, its authors report refusal rates under abliteration that "drop by at most 10%, compared to 70-80% drops in baseline models" (Abu Shairah et al. 2025). One caution from the same field: Qi, Wei, Carlini and co-authors (2024) warn that evaluating such defences "is exceedingly difficult and can easily mislead audiences into thinking that safeguards are more durable than they really are".
The landscape, briefly
Two download counters give the scale, each with its window. The abliterated Qwen3.8-27B GGUF repository of huihui-ai, one publisher of such copies, stood at 2,030,881 downloads in Hugging Face's 30-day count, read at 2026-10-02T20:32:45Z, and 4,019,531 since its creation on 2026-08-16. The upload this page reads, orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF, stood at 191,956 in the 30-day count at the same read, and 282,536 since its creation on 2026-08-26, with 486 likes. Neither counter counts people: Hugging Face counts as a download every request for any GGUF file, even one that only checks the file, and every compressed copy in the upload is split into two to five files, so one person fetching one copy counts at least twice. We ran no count of how many such repositories exist, so this page makes no claim about that.
One word covers all three mechanisms. A repository called "uncensored" or "abliterated" may be a retrain, a weight edit, or a run-time file, and the card is the only place that says which. The section "Read the label", further down, is a short guide to reading one.
One upload, read by its hashes
The verdict first, and it is in the uploader's favour. The uploader is OrcaRouter, which describes itself as "The AI Gateway built for Privacy & Security". As far as public evidence reaches, this upload is what its card says it is: standard llama.cpp quantizations of the uploader's own abliterated full-precision copy of Qwen's model, made with the weight edit of (b); not a fine-tune, not a merge, not a run-time vector. Here is what we can show, in the order we found it out.
Hugging Face publishes the size and the sha256 of every large file in a repository, whether or not you can download it. So without fetching a byte we laid the uploader's full-precision copy (131 shards, 360,000,192,888 bytes in all) beside Qwen's own at its pinned revision (131 shards, the same byte total): 80 of the 131 shards are byte for byte Qwen's; 51 differ, at identical sizes.
Qwen's own tensor index (a tensor is one named block of the model's numbers; the matrices below are tensors) then says which parts of the model live in which shard; the uploader's copy of that index is the same file, since of the 13 files beside the weights only the card differs. Every one of the 51 that changed holds at least one of the 149 matrices the card says it edited, each of them a matrix that writes into the running state, and none of the 80 unchanged files holds one. In all 131 files the evidence agrees with the card.
Then the limit, said plainly: a hash sees a whole file. 1,336 other tensors share those 51 files (the routers, the norms, the attention inputs, the whole vision tower, the part that reads images), and whether they were left alone is the card's word, not something we checked. And a hash says that a file changed, not how: the 80 unchanged files rule out a full fine-tune or a whole-model merge, which would have changed them too, but nothing here tells the card's training-free edit from any other change confined to those same matrices; that part is the card's word as well. What we can prove untouched is 173 tensors, 72.6 % of the model's file bytes.
The compressed copies come next, and here we are fairly confident rather than certain that they are standard quants of this model; which of the two full-precision copies they were made from, the evidence below cannot say. One 54,400,261,312-byte file (54.4 GB) of the upload, the second shard of both its Q6_K and its Q8_0 builds, is byte for byte the same file two other quantizers (publishers of compressed copies) published from the unmodified model (by its byte count, the model's n-gram table, a large lookup table, in its 8-bit form). Eight of the sixteen quant types in the upload sit within 704 bytes of other quantizers' default quants of the unmodified model, which is the signature of the standard process. And the chat template, the text wrapped around every conversation, is Qwen's byte for byte in the one file Hugging Face parses. None of the three tells a quant of the edited parent from a quant of Qwen's own weights: the shared shard and the template are parts the edit leaves alone, and a file's size follows how it was compressed, not the numbers inside it (one third party's two copies at the same level, Q3_K_S, one made from the edited parent and one from Qwen's own model, differ in size by 512 bytes). That the compressed files carry the edit is the card's word, and nothing we found contradicts it.
Now the thing the tree leaves out. The upload's repository metadata (its card's base-model field) names Qwen as the direct base, though the card's own text names the edited parent, so Hugging Face's own tree view of the model skips a generation and leaves the parent out. The parent, in turn, calls itself a fine-tune of a model it edited without training; Hugging Face's metadata offers four labels for a derived model (adapter, merge, quantized, finetune) and none for an edit made without training, so "finetune" is the nearest box. A reader who wants the parent's name has to read the prose.
The uploader's own figures, printed as their unreproduced claims, from the parent card at revision e096800036ec. The GGUF card at e43d00f4e2b8 summarises them as staying "within ±2 pts of the base across MMLU-Pro / GSM8K / CMMLU-style checks", and that list is quoted whole: the one row outside two percentage points is the one its summary leaves out, though the parent card prints that row with the rest. The four are standard tests: MMLU and its harder successor MMLU-Pro ask multiple-choice questions across many school and professional subjects, GSM8K is grade-school maths word problems, and CMMLU is the same kind of knowledge test in Chinese; "0-shot" means the model saw no worked examples first, and "chain of thought" means it reasoned step by step before answering.
The uploader's own scores for its edit of Qwen3.8-Flash-Next: % of questions answered correctly, before and after, with thinking off (unreproduced).
| Test | % before | % after | Questions |
|---|---|---|---|
| MMLU, 0-shot | 90.0 | 87.7 | 300 |
| MMLU-Pro, chain of thought | 77.8 | 76.2 | 400 |
| GSM8K, chain of thought | 92.0 | 93.3 | 150 |
| CMMLU, 0-shot | 81.8 | 81.6 | 500 |
The MMLU row reads as a 2.3-point drop on 300 questions (90.0 → 87.7, the uploader's figures: by our arithmetic, 270 right before and 263 after, a net 7 questions); whether that is noise depends on how many individual questions changed answers, which the uploader does not publish; on its own the figure can be neither trusted nor dismissed. By the parent card's own account, all of these figures were measured on the full-precision model, with thinking off, which is not the model's default; the GGUF card says its quants "inherit these behaviours" but shows no measurement of them, so the compression's own cost is unmeasured by the uploader.
The arithmetic behind the MMLU reading, for readers who want it: one score of about 90 % on 300 questions has a standard error of about 1.7 points; the difference between two such scores taken independently about 2.6; but these two scores come from the same 300 questions, so the right test is a paired one, and its uncertainty can be as little as 0.9 points if every changed answer went the same way, and at most about 2.7 at the most flips the two totals allow.
What an edit like this one costs a model on a fair, matched test, the same questions put to an edited copy and to its unmodified original, is a separate page, designed and not yet run as of 2026-10-05 (UTC); its first pair is an edited and an unmodified copy of a smaller Qwen model, not this upload.
A speed setting, read by its hash
The twist, with the engine's side first. Despite the similar name, Strata (github.com/Niko1221/Strata) is an unrelated project; strata2signal has no affiliation with it. That project is an engine, a program that runs this same model, Qwen3.8-Flash-Next, on one graphics card and the rest of an ordinary PC. It ships an optional "experimental speed projection", off unless the user chooses it. The setup question that offers it says, in the engine's own words, that "its package describes it as a refusal-direction projection (the model declines far fewer requests)". The explanatory note in the file's folder (its README) adds "you are responsible for what the model writes with it on" and "It is not an optimization in the engine", and the engine's README names the same licence the vector's own card names, the Qwen Community License 1.0, though the licence text does not ship beside the file. The file is 483,520 bytes, and its sha256 is the sha256 of a refusal-direction control vector published on Hugging Face on 2026-09-17 (repository Cudecnik/Qwen3.8-Flash-Next-refusal-projection at revision 1886570b24da; its bytes and sha256 are in the kit's pinned-revision list), which that repository lists under two names, -refusal-projection and -refusal-additive, at one hash. By the vector maker's card it is technique (c), the mute on the channel, in the remove flavour. Mainline llama.cpp's control vectors only add, so applying it as a removal needs code that mainline does not have, and the "-additive" name is no contradiction: a vector file holds only the vectors, and the engine that loads it decides what to do with them, add or remove.
Then the part we could not explain: the name says speed. An outside pull request to the engine (a proposed change to its code, #15, opened 2026-09-26) linked that original by name. On 2026-09-27 at 14:47Z the maintainer's own commit, a change saved to the engine's code, shipped the same bytes under the new name, and the pull request was closed without being merged at 15:43Z that day. The README that commit added says that the phrase "refusal-direction projection" comes from "the package's own documentation" and that "the published original uses filenames containing 'refusal'; the bytes are the same": it acknowledges the original, but gives no name and no link. We found no stated reason for the name. We guess no motive, and the pull request's author is not named here.
Read the label, and if you run these files
Reading a label, not a recipe; if you will never download a model yourself, you can skip to When the model is the only book on the shelf. What the card says tells you which of the three kinds you have, and each kind has one thing worth checking.
| What the card says | Which kind | What to check |
|---|---|---|
| The name says "abliterated" | Most likely a weight edit (b); the card says for certain | Compare its full-precision files' hashes with the base model's, as this page did |
| The card names a training dataset | A retrain (a) | Read that dataset's card: a retrain is a different model |
| A small separate vector file, or a switch in the engine | Run-time steering (c) | The switch is yours: read what the engine says it does |
| "Uncensored" with only a system prompt | None of the three | No weights changed; a system prompt is a prompt |
Two limits on those checks. For a weight edit, the changed files should be the ones holding matrices that write into the running state, but the comparison works only on a model split into many files whose upload keeps the original's split: in a dense model (one with no experts, so every part works on every word-piece) every file holds such matrices, so every file changes, and some uploads re-split their files, so nothing lines up. Nor does it work on compressed copies: two quantizers' copies of the same unmodified model at the same level can be different files (lmstudio-community's and unsloth's Q8_0 of Flash-Next differ in size by 83,456 bytes), so a compressed file's hash proves something only where it matches another file exactly, as the upload's one shared shard does. And a small adapter file (a LoRA) is a retrain kept in a separate file; Hugging Face's metadata labels even this page's vector an "adapter", so read the card, not the file size.
If you run these files, general advice with this upload as the example, every hedge kept as the research has it.
- Pin a revision. Both of the uploader's cards changed on 2026-10-02: the parent's, orcarouter/Qwen3.8-Flash-Next-Uncensored at e096800036ec, at 04:18:48Z; the GGUF's, orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF at e43d00f4e2b8, at 04:19:04Z. A card is not frozen; a revision is. If a copy is for a machine that will be offline, keep each file's published sha256 beside it before you disconnect: offline, that is the only fingerprint left to check a copy against.
- Check that the licence file travels. The tag says Apache-2.0 and the card says "inherited from Qwen", but the base model is under the Qwen Community License 1.0. The same Apache tag sits on ISTA-DASLab's widely used Flash-Next quants, and none of the ten Flash-Next GGUF repositories we checked, the llama.cpp organisation's own included, ships Qwen's licence file; the full-precision parent does ship it, byte for byte, under the Apache tag. Mislabelled, in other words, and a shared mislabel. Whether the licence's commercial-use condition binds a free upload is a question for Qwen and the uploader, not one this page can settle.
- Re-download when a fix is announced, and credit the fix. On 2026-09-10 the uploader re-uploaded the first shard of every quant then present with a note that the first shards had carried an all-zero attention setting and that the fix was metadata-only, weights unchanged; the file sizes did not change, which fits. The setting belongs to the model's sparse-attention layers, and current llama.cpp (we read its source) treats a zero there as no setting at all and runs those layers as ordinary full attention, so copies downloaded before that date probably ran them differently from Qwen's design; whether the llama.cpp of early September did the same is unknown, and this page has no measurement of what it changed in the replies. For one week, 2026-09-03 to 2026-09-10, the Q8_0 files were a build that, by a report on the repository's own discussion page, mainline llama.cpp could not load; a plain build replaced it.
- Put the server behind something. The card's example server command listens on every network interface with no key, so anyone who can reach that machine over a network can use the model. Whatever you run it on, give it a door with a lock.
- Know what the gate is. Both repositories use an access gate, Hugging Face's standard option in its default automatic mode, which shares the downloader's username and email address with the uploader. We did not accept the gate, and this page did not need to.
When the model is the only book on the shelf
The last question this page takes up is a practical one: what about a model kept for emergencies, in an offline, "prepper" kind of scenario, where it is basically the only source of knowledge? No network, no second opinion; a laptop, a model on it, and whatever else is in the room. (The model this page reads is a big one: the smallest compressed copy in the upload is 74.8 GB, so the laptop in this picture may well carry a smaller model, and nothing on this page measures one.) Another of our pages has measured something close: a mini PC with no graphics card ran a 26-billion-parameter model, compressed to 4 bits and holding about 18 GB of its 32 GB of memory, at about twelve word-pieces a second on a consumer battery backup, for a projected two hours of continuous answering.
That is the setting where the case for a less-refusing model is at its strongest. If the only book on the shelf declines to say how long to boil water before it is safe, or how to splint a forearm with what is in a car, or how to keep warm overnight in a stalled vehicle, because the question rhymed with something it was trained to decline, the refusal has a real cost and nothing else is there to pay it. Over-refusal is not a laboratory curiosity there; it is the book closing when you need it open. How often that happens is a measurement, not a given, and we found no published figure for survival questions. The nearest figure for the upload this page reads is the uploader's own, unreproduced, on XSTest's 250 harmless prompts, a set built to catch over-refusal in general rather than on survival questions: the model refused 0.4 % of them in its default mode, with thinking on, both before and after the edit; the drop they report, 9.6 % to 1.2 %, is with thinking switched off.
What a fair measurement of this would look like: harmless, everyday survival questions, first aid, water, food, shelter, put to an edited model and to its unmodified original, counting refusals only, which would show how often each copy answers and nothing about whether the answers are right. The one public set of reference answers we found for questions like these is synthetic (written by a model) and in free text, so no answer would be graded for correctness against it, and none against our own guesses either. The companion bench test, the separate page named above, includes this count in its plan as an optional extra beside its main measurement, what an edit costs in capability; it is not promised. Nothing on that list is a dangerous-capability question, and nothing on it is a recipe.
But the same setting is where a confident error costs the most, and nothing on this page shows that removing refusals makes a model more correct. A model that has had its no removed has not had its guessing removed. A confidently wrong answer about which wild plant is safe to eat can cost more than a refusal.
So the robust design, the one we would build for ourselves, is not the model alone: it is a model paired with offline references it can cite, an offline encyclopedia and a medical reference beside it, so that every answer that matters carries a page you can open and check with your own eyes. That works only when the program looks the passage up in the offline library and shows it to you beside the answer (what the field calls retrieval-augmented generation, after Lewis et al. 2020); a model asked to cite from memory can invent a source as confidently as it invents a fact. Even then, a cited page can show you what the reference says, not whether the plant in your hand is the one on the page. The model is the librarian, not the library. One such door, named in its own words and not as a recommendation: Kiwix describes itself as "a nonprofit organisation making free knowledge accessible where the Internet is not", whose tools "started with offline access to Wikipedia"; its catalogue lists a "WikiMed Medical Encyclopedia", described there as "The largest medical encyclopedia, from Wikipedia", in English and other languages. We have not tested it.
For the model this page reads, in its default mode, the edit buys little on harmless prompts by the uploader's own figures, and nothing on this page shows what it costs on the questions an offline reader would ask; on this page's evidence, the choice that matters most offline is the references beside the model, not which copy of it you keep. And have the first-aid basics on paper too: a printed manual needs no battery, and it cannot invent a page.
How this workshop handles these files
This workshop keeps a few rules for files like these. Any bench test it runs on such a file uses harmless prompts only, and no such file serves anyone outside the workshop. Flash-Next itself is used here internally only; serving it to anyone outside would need a closer reading of its licence first. Three habits apply to anyone who runs one of these models, because once its refusals are gone, what is left of the no is the law and the person running it: pin a revision, read the licence, and know which of the three kinds you have (the card says which, if you read it). For a machine with no network, the answer is the one in the section above: the model beside the books, not instead of them.
What to take with you
- "Uncensored" names a result, not a method: a model changed to decline far less often, by a retrain, by one permanent edit to its weights, or by a switch while it runs, and only the card says which. What goes includes the serious refusals: by the uploader's own figures, the edit behind the upload this page reads takes out most of the model's refusals of harmful requests.
- 80 of the 131 weight files in the uploader's full-precision copy are byte for byte Qwen's own, and all 51 that changed hold a part the card says it edited, by the sha256 Hugging Face publishes for every large file at the two pinned revisions: at file level, the evidence agrees with the card. The 1,336 other tensors inside the changed files are the card's word; 173 tensors, 72.6 % of the model's file bytes, are provably untouched.
- The "experimental speed projection" in the Strata engine (github.com/Niko1221/Strata) is, byte for byte, a published refusal-direction vector under two names: 483,520 bytes, one sha256, three names counting the engine's. It is optional and off unless the user chooses it, and the engine's setup question says what it does; its name does not.
- The edit changes the weights, never the licence. Qwen's licence has no use policy; its two conditions are the notice and commercial use by businesses that host models or sell AI coding and office assistants, and Qwen's licence file does not travel with any of the ten Flash-Next GGUF repositories we checked.
- A weight edit takes out a habit; it adds no knowledge. The uploader's own scores for this edit move within about two percentage points either way (one 2.3-point drop on 300 questions), and by the uploader's figures, in its default mode the model refused 0.4 % of 250 harmless test prompts both before and after; for a model that is your only source, the robust set-up is the model beside offline references it can cite.
How to check our work — and see it live
- Re-run the census with no code. Open both repositories' file lists at the pinned revisions (Qwen's and the uploader's full-precision copy), where each file is listed with its size; the sha256 of every file is published in each repository's Hugging Face API record at that revision, which a browser opens without an account (the uploader's own file pages ask for one); compare the 131 pairs, then look up in Qwen's tensor index which parts each file holds: every changed file holds at least one of the 149 matrices the card names, and no unchanged file holds one.
- Check the shared shard. The 54,400,261,312-byte file is listed, with its sha256, in lmstudio-community's quant of the unmodified model, in bartowski's, and, for the upload, in its API record at that revision (its file list shows the file's name and size).
- Check the vector. Its repository's Hugging Face API record at revision 1886570b24da publishes the file's size and sha256 under both names, and the kit's pinned-revision list carries both; the engine's copy sits in its repository under data/experimental-speed-projection, and the pull request and commit are numbered and dated above. We name both repositories as text and link neither one's files: the engine folder's own README carries settings this page does not reproduce.
- Read the licence. Qwen Community License 1.0, at the pinned revision; and the two use policies named above, Meta's and Google's, to see what one looks like.
- Take the kit. Four files and nothing else: the shard census (the 80 and the 51, by file name), the quant size matches (sixteen rows), the pinned-revision list with every file's bytes and sha256 that a claim on this page rests on, and the licence table for thirteen repositories (tag, revision, licence file present). No model card, no README, no code. Open the kit.
- Read the papers. Every one is linked at arXiv, the free online archive of research papers, where it is cited, in its own words.
We tested every link in this article that leads off it, without an account, on 2026-10-05 between 01:26:56Z and 01:27:02Z: 48 of the 49 opened, and the 49th, this page's own kit, opens with this page.
The rest of the seminar
- A short history of Ollama and A short history of vLLM: the two model servers this workshop benches with.
- The compressed photograph: what codes like Q4_K_M on a GGUF mean.
- Reading the answer instead of writing it: the six-way test this workshop scores models with, the one the companion bench test is designed to use.
- The licence ledger: every model our products run, and every model we benched on the way, with its licence read first-hand.
- Planned, no link and no date: the companion bench page (what a refusal-removing edit costs a model, on a fair, matched test); the Strata engine bench page (github.com/Niko1221/Strata).
If something here is wrong, or there is a check you want next, say so.
Who ran this, and thanks
Thanks to Andy Arditi and co-authors, whose 2024 paper is the primary record of the direction this page is about, and to the authors of every paper linked above, who publish their abstracts where anyone can read them. Thanks also to Qwen for publishing its model's tensor index and licence beside its weights, which is what made the census possible without a download; to Hugging Face for publishing the size and sha256 of every large file in a repository, gated or not; to the uploader, OrcaRouter, for a dated fix note and for publishing its own evaluation figures, which this page prints as theirs; to the author of the published vector for a card that names the base model's licence; to the maintainer of the Strata engine (github.com/Niko1221/Strata) for a setup question that says in plain words what its optional file does; and to Eric Hartford, whose 2023 article is still the clearest account of where the accidental no came from. A small human team asked for this page, chose what it would and would not claim, and signed off on it; a fleet of AI agents read every source listed below and drafted it under that team's rulings.
Sources
All read on 2026-10-02 (UTC) unless a line says otherwise; each by the small team's own agents, at the time given, with receipts kept in the workshop's archive.
- Hugging Face API records with file hashes (?blobs=true) for Qwen/Qwen3.8-Flash-Next at de4b8e4d43b9 and orcarouter/Qwen3.8-Flash-Next-Uncensored at e096800036ec (19:30Z; re-read 20:27:39Z and 20:59:20Z); orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF at e43d00f4e2b8 (19:27:51Z); Qwen's model.safetensors.index.json at de4b8e4d (19:34Z; re-read 20:27:59Z and 20:59:20Z); the API records of bartowski's, lmstudio-community's, mradermacher's, unsloth's, ggml-org's and ISTA-DASLab's Flash-Next GGUF repositories, and of Cudecnik/Qwen3.8-Flash-Next-refusal-projection at 1886570b24da (19:31-19:38Z; the thirteen licence rows re-read 20:28:28Z).
- Hugging Face's download and like counters for the two repositories named in the landscape section, with both counters expanded (20:32:45Z, re-read 21:00:10Z), and the huggingface_hub client's definition of each counter at its d711944d commit, lines 804-807.
- The uploader's two model cards at the revisions named (19:27-19:30Z): the evaluation table and the thinking-off / thinking-on over-refusal pair from the parent card; the summary lines from the GGUF card. The repository's tree at each of its seven earlier revisions and its discussion pages, for the fix of 2026-09-10 and the Q8_0 swap (19:34-19:37Z).
- llama.cpp at commit bed0a856 (19:34-19:41Z and 20:29:38Z): the handling of a zero attention setting in its qwen4exp model code, and the control-vector code, which adds.
- The repository of the Strata engine (github.com/Niko1221/Strata) at commit 8fc40dde (20:29:33Z): setup.py lines 1689-1692 and the engine's README.md lines 106, 139-140 and 218-219; the folder README of data/experimental-speed-projection (20:17Z); GitHub's record of pull request #15 and of the commits touching that folder (20:28:59-20:29:13Z; re-read 21:03:25Z).
- arXiv's export API for every paper linked (20:17:22Z, 20:29:14-20:29:30Z, 21:00:31Z and, for Wang et al. 2024, 23:23:26Z; for Qi, Wei, Carlini and co-authors 2024 and Lewis et al. 2020, 2026-10-04 at 23:19:08Z): titles, authors, dates and abstracts, which are quoted as they stand.
- Eric Hartford, "Uncensored Models", erichartford.com, published 2023-05-15 (the page reads "Updated May 22, 2023"), read 20:28:53Z, 20:30:17Z and between 22:51Z and 22:57Z; its prose is quoted and nothing else from it is reproduced.
- The Llama 2 paper's Section 4 headings (ar5iv, between 22:51Z and 22:57Z); Meta's "Llama 3 Acceptable Use Policy" and Google's "Gemma Prohibited Use Policy" (between 22:51Z and 22:57Z), read for their existence and titles only.
- Hugging Face's documentation of gated repositories and of the base_model_relation field (20:31:45Z), and of how downloads are counted (read 2026-10-04, 23:19:20Z).
- Qwen's model card, its configuration file (the four parallel streams of the running state in D1's caption) and the Qwen Community License 1.0 at revision de4b8e4d (read 2026-09-29 and 2026-10-02 19:30-19:32Z; the licence's git blob matched against the uploader's parent repository at 20:28Z).
- Our own published page on a mini PC on battery (Two Hours at 12 tok/s — On Battery, exhibit thirty-two, published 2026-08-31), for the offline section's memory, speed and battery figures (read 2026-10-04, 23:16-23:19Z).
- Kiwix's own home page, for its self-description, and its catalogue's machine-readable listing, for the WikiMed entries and their one-line descriptions (read 2026-10-03, 00:02-00:07Z); nothing downloaded.
Corrections and later measurements will be added below, each dated (UTC), with a window at both ends where one applies and saying in plain words what it counts.