You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Compatible checkpoint ONLY (read this first)

These weights are a Mia single-DGX-Spark recipe splice. They are NOT a drop-in ablit for other Qwen3.8-Flash-Next NVFP4 quants.

You may use this repo only with Mia's single-Spark Flash Next checkpoint layout:

Do not copy these shards, or this o_proj edit, onto RadixArk / NVIDIA ModelOpt experts-only NVFP4, Inferact, official BF16/FP8, GGUF, MLX, or any other NVFP4. Attention o_proj is MXFP8 in this layout and BF16 in the NVIDIA/RadixArk experts-only recipe. Mixing them will not load correctly and is not this ablit.

Responsible Use Agreement

This model has had safety refusals removed. That makes it useful for red-teaming, security research, evaluation, and unfiltered assistant tasks — and also removes guardrails a user must therefore supply themselves.

Prohibited uses (you must agree before access is granted):

  • Anything involving the sexual exploitation or endangerment of minors.
  • You must be of age 18 years or older to use and download this model.
  • You agree any information generated that can cause harm in terms of generating recipe, knowledge to make any materials/substances is your own input and responsibility. You will be accountable for any harm/damage caused by your action/input.
  • Content promoting self-harm or suicide.
  • Generation of material that is illegal in your jurisdiction, or that targets real individuals for harassment, doxxing, or fraud.
  • Any use prohibited by the upstream Qwen Community License.

You are responsible for adding appropriate safety filtering, human review, and access controls for your deployment. The weights are provided as-is, with no warranty. The license is inherited from the upstream Qwen base model — review and comply with it before use or redistribution.

Log in or Sign Up to review the conditions and access this model content.

keys-Qwen3.8-flash-next-ablit-Mia-Single-Spark-only

STOP — Mia single-Spark NVFP4 only

This ablit cannot be used on other models. It is valid only on Mia's single-DGX-Spark Qwen3.8-Flash-Next NVFP4 checkpoint:

Do not implement, copy, or “port” these shards onto other NVFP4 quants (NVIDIA ModelOpt experts-only / RadixArk, Inferact, PLE-NVFP4 variants, official BF16/FP8, GGUF, MLX). Those trees store QSA o_proj as BF16 (or another layout). This dest stores the edited writers as MXFP8 in the Mia 34-shard recipe. Dropping these files on another NVFP4 will not apply this ablit and can break load.

Abliterated Qwen3.8-Flash-Next weights for one DGX Spark (TP=1) via the Mia single-Spark vLLM recipe. MTP, PLE n-gram table, routed experts, GDN out_proj, vision, and L0–14 attention writers stay stock.

Compatible base (ONLY) Mia-AiLab/Qwen3.8-Flash-Next-NVFP4 · local-inference-lab/Qwen3.8-Flash-Next-NVFP4
Not compatible Any other Flash Next NVFP4 / FP8 / BF16 / GGUF / MLX
Refusal suite 32/32 BYPASS (QuantTrio-style hard suite) · 0 refuse · 0 garble · 0 errors
Ablit QSA self_attn.o_proj L15, 19, 23, 27, 31, 35, 39, 43, 47 (9 tensors + scales) MXFP8 · MTP stock · experts stock
Runtime Mia single-Spark vLLM (vllm/vllm-openai:qwen38-flash-next) · TP=1 · MTP D=3 · KV_CACHE_DTYPE=fp8 · util ≤ 0.85

Upstream: Qwen/Qwen3.8-Flash-Next · serve recipe: MiaAI-Lab


⚠️ Compatible checkpoint (mandatory)

This is a splice into the Mia single-Spark recipe layout (34 safetensors shards, MXFP8 attention, NVFP4 routed experts, NVFP4 packed PLE).

Other Flash Next NVFP4 Why this dest does not apply
NVIDIA ModelOpt / RadixArk experts-only NVFP4 QSA o_proj is BF16, not MXFP8; ~206 shards; FP8 PLE — not this ablit
Any other vendor NVFP4 of Flash Next Different shard map / o_proj dtype
Inferact / other PLE-NVFP4 trees Different shard map / o_proj dtype
Official BF16 or FP8 Different quant entirely
GGUF / MLX Different format
nvidia/Qwen3-8B-NVFP4 or other Qwen3 models Different model

If you need the same kind of edit on another quant, re-encode the same 9 QSA layers into that tree’s o_proj format. Do not copy these 34 shards.


⚠️ Responsible Use & gated access (required)

WARNING: This model has had safety refusals removed. That makes it useful for red-teaming, security research, evaluation, and unfiltered assistant tasks — and also removes guardrails you must supply yourself.

Access is gated. By requesting Hugging Face access, downloading, or using these weights, you agree to the terms in RESPONSIBLE_USE.md.


Abliteration

Method Residual-writer splice: QSA self_attn.o_proj only
Layers 15, 19, 23, 27, 31, 35, 39, 43, 47 (12-way QSA grid; L3/L7/L11 stock)
GDN out_proj stock (requant of donor GDN writers collapsed to stock MXFP8)
MTP stock (model-00034-of-00034.safetensors hardlinked)
Routed experts stock
PLE n-gram stock
Chat template stock (not a third-party template)
Format Mia recipe MXFP8 (F8_E4M3 + e8m0 scales, group 32)
Meta ABLIT_META.json

Speculative decoding (in-checkpoint MTP, D=3) is kept target-aligned: do not apply this edit and then swap in another NVFP4’s MTP/PLE.


Files of interest

Path Purpose
model-*-of-00034.safetensors + index Full Mia-layout checkpoint (MTP stock)
COMPATIBILITY.md Hard lock: Mia single-Spark NVFP4 only
ABLIT_META.json Edit fingerprint / which tensors changed
results/refusal32-summary.json 32/32 suite labels (no completions)
hf_quant_config.json MXFP8 attention + NVFP4 experts

Download

# after access is approved
hf download drowzeys/keys-Qwen3.8-flash-next-ablit-Mia-Single-Spark-only \
  --local-dir ~/models/Qwen3.8-Flash-Next-Abliterated-NVFP4-spark-oproj-L15-47

Serve only with the Mia single-Spark recipe (TP=1, PLE packed table, MTP 3). Point MODEL_DIR at the download. Do not load this tree as a generic “NVFP4 Flash Next” in SGLang/vLLM recipes written for RadixArk.


Benches (one DGX Spark / GB10, this dest)

Refusal: 32/32 BYPASS, 0 refuse, 0 garble.
Warm decode vs stock Mia NVFP4 on this host (MTP D=3, thinking off): prose ~21–22 tok/s (AL ~1.0); code ~31–35 tok/s (AL ~2.0). Stock code AL on this host was ~2.2 — remeasure AL if you change D.


Credits

Base model

https://huggingface.co/Qwen/Qwen3.8-Flash-Next

Downloads last month
1,160
Safetensors
Model size
93B params
Tensor type
BF16
·
F8_E4M3
·
U8
·
I64
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for drowzeys/keys-Qwen3.8-flash-next-ablit-Mia-Single-Spark-only

Quantized
(254)
this model