Instructions to use ProCreations/MiMo-V2.6-Flash-RL-IQ3_XXS-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use ProCreations/MiMo-V2.6-Flash-RL-IQ3_XXS-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf ProCreations/MiMo-V2.6-Flash-RL-IQ3_XXS-GGUF:IQ3_XXS # Run inference directly in the terminal: llama cli -hf ProCreations/MiMo-V2.6-Flash-RL-IQ3_XXS-GGUF:IQ3_XXS
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf ProCreations/MiMo-V2.6-Flash-RL-IQ3_XXS-GGUF:IQ3_XXS # Run inference directly in the terminal: llama cli -hf ProCreations/MiMo-V2.6-Flash-RL-IQ3_XXS-GGUF:IQ3_XXS
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf ProCreations/MiMo-V2.6-Flash-RL-IQ3_XXS-GGUF:IQ3_XXS # Run inference directly in the terminal: ./llama-cli -hf ProCreations/MiMo-V2.6-Flash-RL-IQ3_XXS-GGUF:IQ3_XXS
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf ProCreations/MiMo-V2.6-Flash-RL-IQ3_XXS-GGUF:IQ3_XXS # Run inference directly in the terminal: ./build/bin/llama-cli -hf ProCreations/MiMo-V2.6-Flash-RL-IQ3_XXS-GGUF:IQ3_XXS
Use Docker
docker model run hf.co/ProCreations/MiMo-V2.6-Flash-RL-IQ3_XXS-GGUF:IQ3_XXS
- LM Studio
- Jan
- vLLM
How to use ProCreations/MiMo-V2.6-Flash-RL-IQ3_XXS-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ProCreations/MiMo-V2.6-Flash-RL-IQ3_XXS-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ProCreations/MiMo-V2.6-Flash-RL-IQ3_XXS-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ProCreations/MiMo-V2.6-Flash-RL-IQ3_XXS-GGUF:IQ3_XXS
- Ollama
How to use ProCreations/MiMo-V2.6-Flash-RL-IQ3_XXS-GGUF with Ollama:
ollama run hf.co/ProCreations/MiMo-V2.6-Flash-RL-IQ3_XXS-GGUF:IQ3_XXS
- Unsloth Desktop
- Pi
How to use ProCreations/MiMo-V2.6-Flash-RL-IQ3_XXS-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ProCreations/MiMo-V2.6-Flash-RL-IQ3_XXS-GGUF:IQ3_XXS
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ProCreations/MiMo-V2.6-Flash-RL-IQ3_XXS-GGUF:IQ3_XXS" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use ProCreations/MiMo-V2.6-Flash-RL-IQ3_XXS-GGUF with Docker Model Runner:
docker model run hf.co/ProCreations/MiMo-V2.6-Flash-RL-IQ3_XXS-GGUF:IQ3_XXS
- Lemonade
How to use ProCreations/MiMo-V2.6-Flash-RL-IQ3_XXS-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull ProCreations/MiMo-V2.6-Flash-RL-IQ3_XXS-GGUF:IQ3_XXS
Run and chat with the model
lemonade run user.MiMo-V2.6-Flash-RL-IQ3_XXS-GGUF-IQ3_XXS
List all available models
lemonade list
- Hermes Agent
How to use ProCreations/MiMo-V2.6-Flash-RL-IQ3_XXS-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ProCreations/MiMo-V2.6-Flash-RL-IQ3_XXS-GGUF:IQ3_XXS
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ProCreations/MiMo-V2.6-Flash-RL-IQ3_XXS-GGUF:IQ3_XXS
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use ProCreations/MiMo-V2.6-Flash-RL-IQ3_XXS-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ProCreations/MiMo-V2.6-Flash-RL-IQ3_XXS-GGUF:IQ3_XXS
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ProCreations/MiMo-V2.6-Flash-RL-IQ3_XXS-GGUF:IQ3_XXS" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Configure Hermes
# Install Hermes:
curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash
hermes setup# Point Hermes at the local server:
hermes config set model.provider custom
hermes config set model.base_url http://127.0.0.1:8080/v1
hermes config set model.default ProCreations/MiMo-V2.6-Flash-RL-IQ3_XXS-GGUF:IQ3_XXSRun Hermes
hermesMiMo-V2.6-Flash-RL: calibrated IQ3_XXS
This is a lossy, importance-calibrated IQ3_XXS conversion of the original MXFP4 routed expert weights. It prioritizes retention: all non-expert tensors keep their original decoded values, using FP32 or exactly representable BF16 storage. Calibration and evaluation use the original released checkpoint as the reference, not another 3-bit model or a dequantized NVFP4 approximation.
Source: XiaomiMiMo/MiMo-V2.6-Flash-RL,
revision 3b38d063180c3e4aed9691fdc735f3d10b266ee4, MIT license.
For exact preservation of all released weight values, see the separately
verified NVFP4 conversion.
Calibration
The importance matrix is collected from the original MXFP4 experts with cuBLAS F32 storage/accumulation and BF16 KV, bypassing the fast runtime's FP4/Q8 activation quantizers. The cuBLAS handle permits TF32 arithmetic for eligible large multiplications, so this is not a full-mantissa FP32 calibration claim. The selected corpus contains 1,051,701 base tokens plus 167,691 extension tokens. After serialization and runtime tokenization, collection processes 257 chunks at 4,096 context and 10 chunks at 16,384 context: 1,216,512 tokens. All 141 expert projection count sums independently confirm eight routes per processed token. The extension adds complete long code and math examples and deterministic tool/JSON transcripts with checked answers.
Public source datasets are OpenCodeReasoning, OpenR1-Math-220k, ultrachat_200k, and alpaca-gpt4-chinese. Their immutable revisions, selection seeds, complete-example filters and hashes are recorded in the manifests. Question-hash splits and global deduplication separate calibration from the held-out code, math, English and Chinese evaluation text. Executable coding checks are separate handcrafted tasks, not calibration examples.
This is an explicitly mixed-type release: 129 of 141 routed-expert projection
matrices use IQ3_XXS. The 12 matrices in layers 1, 7, 11 and 25 retain their
original MXFP4 representation because at least one expert did not meet the
256-observation coverage floor. All 649 model tensors passed the shape/type and
preservation audit. The importance matrix is included; its SHA256 is
7a82cf559878c59e7a61075608db3cd7fc0c7677354d9d932be6daec6eaadfb7.
The original MXFP4 weights were dequantized by the quantizer and requantized to
IQ3_XXS using that matrix. This is lossy. The exact NVFP4 transcode was not used
as an intermediate. No uncalibrated routed expert was silently forced to IQ3.
The full recipe and per-expert coverage are in reports/.
Measured retention
The matched four-domain text comparison evaluates 65,504 next-token positions from held-out code, math, English and Chinese text (eight 4,096-token chunks per domain, scoring each chunk's second half). Both models use the same cuBLAS F32 storage/accumulation, TF32-permitted arithmetic and BF16 KV settings.
| Domain | Original PPL | IQ3 PPL | PPL ratio | Top-token agreement |
|---|---|---|---|---|
| Code | 2.3774 | 2.396366 | 1.007978 | 93.442% |
| Math | 1.5172 | 1.527828 | 1.007005 | 96.989% |
| English | 9.0561 | 8.166138 | 0.901728 | 85.778% |
| Chinese | 12.6352 | 10.270224 | 0.812826 | 83.531% |
Lower PPL on English/Chinese does not establish higher task accuracy or identical
behavior; the output distributions changed. In particular, 83.531% Chinese
top-token agreement is not an extremely-high-retention guarantee. This release
is calibrated for retention, but broad downstream task evaluation was not
completed before the cloud conversion job was closed. See
heldout-iq3-strict.json for the numerical
results and the precise scope of its completion status.
These measurements describe this candidate and protocol. Perplexity ratios and top-token agreement are not percentages of general intelligence or task accuracy. The saved llama.cpp baseline logits are clipped and uint16 encoded; perplexity ratios use the original unclipped PPL logs. Runtime kernels, KV dtype, activation quantization and sampling are reported separately from weight quantization.
Loading and sampling
The eight target shards total 141,408,349,312 bytes (131.697 GiB). Keep all eight in the same directory and load the first shard. Download with:
hf download ProCreations/MiMo-V2.6-Flash-RL-IQ3_XXS-GGUF --local-dir mimo-iq3
cd mimo-iq3
bash tools/build_workstation_runtime.sh
The patched runtime targets llama.cpp revision
58367713a6935c0810103378144008df32e3d5db. A reference-style CUDA build bypasses
low-bit activation kernels:
cmake -S llama-mimo-nvfp4 -B llama-mimo-nvfp4/build-reference -G Ninja \
-DCMAKE_BUILD_TYPE=Release -DGGML_CUDA=ON -DGGML_CUDA_FORCE_CUBLAS=ON \
-DCMAKE_CUDA_ARCHITECTURES=120 -DLLAMA_CURL=OFF
cmake --build llama-mimo-nvfp4/build-reference -j 16 --target llama-server
GGML_CUDA_CUBLAS_COMPUTE_TYPE=f32 GGML_CUDA_DISABLE_GRAPHS=1 \
./llama-mimo-nvfp4/build-reference/bin/llama-server \
-m gguf/MiMo-V2.6-Flash-RL-IQ3_XXS-00001-of-00008.gguf \
-ngl 999 --tensor-split 0.47,0.53 -c 4096 -b 512 -ub 512 \
--flash-attn on --cache-type-k bf16 --cache-type-v bf16 \
--temp 1 --top-p 0.95 --top-k 0 --min-p 0 \
--fit off --parallel 1 --host 127.0.0.1 --port 18085
This reference-style recipe assumes enough total GPU memory, such as the two 96GB GPUs used for validation. CPU offload is possible, but single-card speed, long context and speculative decoding are not qualified by this IQ3 release. Changing kernels, KV precision, or TF32 settings changes the numerical profile.
Use temperature 1.0 and top_p 0.95, with top_k=0 and min_p=0 to avoid additional runtime-default truncation. Allow enough output tokens for reasoning and count truncated responses as incomplete. The original DFlash weights are reusable from the NVFP4 repository; their inclusion does not guarantee support in every runtime. This text-oriented GGUF release does not certify multimodal generation.
Reproduction
The tools/ directory includes the pinned runtime patches, original-weight
GGUF conversion, calibration selection/collection, recipe generation, quantizer
invocation, and independent artifact verification. reports/ records upstream
and dataset revisions, corpus hashes, actual calibration consumption, the
quantization recipe, held-out measurements, and every published shard SHA256.
After quantization, 103 original FP32 matrices were stored as BF16 only after
checking that every value was exactly representable. Independent readback
restored their original FP32 bit patterns. All other preserved tensors retain
the original stored bytes. See
gguf-release-verification.json.
Runtime patches and experimental tuning options are supplied for reproduction;
including an option is not a claim that it improves IQ3 accuracy or performance.
- Downloads last month
- 557
3-bit
Model tree for ProCreations/MiMo-V2.6-Flash-RL-IQ3_XXS-GGUF
Base model
XiaomiMiMo/MiMo-V2.6-Flash-RL
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp# Start a local OpenAI-compatible server: llama serve -hf ProCreations/MiMo-V2.6-Flash-RL-IQ3_XXS-GGUF:IQ3_XXS