Instructions to use Hikari07jp/Ternary-Bonsai-2-27B-Abliterated-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Hikari07jp/Ternary-Bonsai-2-27B-Abliterated-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Hikari07jp/Ternary-Bonsai-2-27B-Abliterated-GGUF:Q2_0 # Run inference directly in the terminal: llama cli -hf Hikari07jp/Ternary-Bonsai-2-27B-Abliterated-GGUF:Q2_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Hikari07jp/Ternary-Bonsai-2-27B-Abliterated-GGUF:Q2_0 # Run inference directly in the terminal: llama cli -hf Hikari07jp/Ternary-Bonsai-2-27B-Abliterated-GGUF:Q2_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Hikari07jp/Ternary-Bonsai-2-27B-Abliterated-GGUF:Q2_0 # Run inference directly in the terminal: ./llama-cli -hf Hikari07jp/Ternary-Bonsai-2-27B-Abliterated-GGUF:Q2_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Hikari07jp/Ternary-Bonsai-2-27B-Abliterated-GGUF:Q2_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf Hikari07jp/Ternary-Bonsai-2-27B-Abliterated-GGUF:Q2_0
Use Docker
docker model run hf.co/Hikari07jp/Ternary-Bonsai-2-27B-Abliterated-GGUF:Q2_0
- LM Studio
- Jan
- vLLM
How to use Hikari07jp/Ternary-Bonsai-2-27B-Abliterated-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Hikari07jp/Ternary-Bonsai-2-27B-Abliterated-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Hikari07jp/Ternary-Bonsai-2-27B-Abliterated-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Hikari07jp/Ternary-Bonsai-2-27B-Abliterated-GGUF:Q2_0
- Ollama
How to use Hikari07jp/Ternary-Bonsai-2-27B-Abliterated-GGUF with Ollama:
ollama run hf.co/Hikari07jp/Ternary-Bonsai-2-27B-Abliterated-GGUF:Q2_0
- Unsloth Desktop
- Pi
How to use Hikari07jp/Ternary-Bonsai-2-27B-Abliterated-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Hikari07jp/Ternary-Bonsai-2-27B-Abliterated-GGUF:Q2_0
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Hikari07jp/Ternary-Bonsai-2-27B-Abliterated-GGUF:Q2_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Hikari07jp/Ternary-Bonsai-2-27B-Abliterated-GGUF with Docker Model Runner:
docker model run hf.co/Hikari07jp/Ternary-Bonsai-2-27B-Abliterated-GGUF:Q2_0
- Lemonade
How to use Hikari07jp/Ternary-Bonsai-2-27B-Abliterated-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Hikari07jp/Ternary-Bonsai-2-27B-Abliterated-GGUF:Q2_0
Run and chat with the model
lemonade run user.Ternary-Bonsai-2-27B-Abliterated-GGUF-Q2_0
List all available models
lemonade list
- Hermes Agent
How to use Hikari07jp/Ternary-Bonsai-2-27B-Abliterated-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Hikari07jp/Ternary-Bonsai-2-27B-Abliterated-GGUF:Q2_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Hikari07jp/Ternary-Bonsai-2-27B-Abliterated-GGUF:Q2_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Hikari07jp/Ternary-Bonsai-2-27B-Abliterated-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Hikari07jp/Ternary-Bonsai-2-27B-Abliterated-GGUF:Q2_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Hikari07jp/Ternary-Bonsai-2-27B-Abliterated-GGUF:Q2_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Ternary-Bonsai-2-27B-Abliterated (PQ2_0)
๐ง PREVIEW RELEASE (v0.1) โ feedback wanted. This is an early preview of a quant-native abliteration of the Bonsai 2 pack. The numbers below are from my own small harness, so I would rather have the community poke at it than call it finished. Please open a Discussion in the Community tab with anything you find โ good or bad.
Status: preview (v0.1) โ not a final release. A refusal-reduced build of
prism-ml/Ternary-Bonsai-2-27B-gguf (PQ2_0, 2.13 bpw, 7.21 GB). The pack is edited
in its own deployed quantization format: there is no BF16 dequantization anywhere in
the pipeline and no re-quantization.
| base pack | prism-ml/Ternary-Bonsai-2-27B-gguf (PQ2_0, ggml type 142) |
| upstream base | Qwen/Qwen3.8-27B |
| file size | 7,206,168,928 B โ identical to the parent pack |
| sha256 | 41a362f422b70a8c2dc74a3cc14447ad0ea702c440f0dbe41dc1796da7b7e342 |
| tensors changed | 400 of 851 (2-bit codes only); the other 451 are byte-identical |
| block scales | untouched |
| recommended mode | thinking on, reasoning_effort=medium |
Method (high level)
The official pack's stored 2-bit codes were updated directly, in place, so the shipped artifact is the edited pack: same geometry, same size, same kernel path, no runtime hook, no adapter. A full-precision refusal-free reference of the same base family was used only to define what to change; that change was then written onto the pack's existing quantization lattice with an unbiased code-rounding schedule, so that the expected weight movement equals the intended edit even though each individual code can only move in whole lattice steps. The edit was amplitude-swept and selected on measured behaviour.
Direction extraction, the reference construction and the rounding schedule are intentionally not published here.
Measured behaviour
Strict labeler, greedy decoding, PrismML fork runtime. "refusal" counts both plain refusals and refusal-then-pivot answers; "comply" is substantive compliance.
| instrument | parent pack | this model |
|---|---|---|
| medium ยท harmful n40 refusal | 37/40 | 0/40 |
| medium ยท harmful n40 substantive comply | 3 | 37 |
| medium ยท hardest-stubborn n25 refusal | 25/25 | 0/25 |
| xhigh ยท harmful n40 refusal | 25/40 | 1/40 |
| xhigh ยท harmful n40 substantive comply | 13 | 21 (16 answers empty โ see caveats) |
| medium ยท math / coding / tool / agentic (executed) | 10/10 ยท 19/20 ยท 15/15 ยท 5/5 | 10/10 ยท 19/20 ยท 15/15 ยท 5/5 (no change) |
| medium ยท one-shot GSM8K-30 (greedy) | 28/30 | 27/30 |
| medium ยท one-shot SST-2-30 (greedy) | 29/30 | 29/30 |
| medium ยท harmless n20 | 11 substantive, 7 shallow | 14 substantive, 5 shallow, 1 empty |
Capability is unchanged within measurement noise on every probe: the executed math/coding/tool/agentic battery is identical item-for-item, GSM8K is 27 vs 28 of 30 and SST-2 is 29 vs 29 of 30 (greedy). The only behavioural difference the harness sees is the one that was intended: on the same harmful probe the parent refuses with a pivot while this build answers substantively.
Feedback I am looking for
- Does the edit hold up on other prompt sets / languages / framings you care about, or does refusal come back?
- Regressions: coding, math, long-context, tool use, multi-turn agentic behaviour โ anything that looks broken or degraded compared to the parent pack.
- xhigh vs medium: medium is the recommended mode here; reports about xhigh (the template default) are especially useful, including instances of empty answers.
- Anything that looks like lattice damage: repeated text, language mixing, degraded fluency.
- If it works for you: your hardware, runtime/flags and sampling settings, so the preview notes can carry a compatibility list.
Reproduce-it-yourself numbers are more useful than impressions โ the labeler I used, the instruments, and the exact prompts are all things I can share in the discussion.
Usage
The chat template defaults to reasoning_effort=xhigh; this build is tuned and measured
at medium, so ask for it explicitly:
llama-server -m Ternary-Bonsai-2-27B-Abliterated-PQ2_0.gguf \
-ngl 999 -fa on -c 32768 --jinja --temp 1.0 --top-p 0.95 --top-k 20 \
--chat-template-kwargs '{"reasoning_effort": "medium"}'
Any OpenAI-compatible client works against http://host:port/v1/chat/completions.
Caveats
- Safety-reduced model. Refusal is largely removed in medium mode. Intended for research and uncensored local use.
- xhigh (the template default) is the weaker mode. Refusal is still down to 1/40
there, but 16/40 answers came back empty because the long xhigh thinking exhausted the
2500-token budget in our harness; raise
max_tokens, or use medium. - The capability battery is small (10 math, 20 coding, 15 tool, 5 two-step agentic) and this is a single-turn gate; multi-turn robustness was not measured.
- Only the 400 edited tensors differ from the parent; integrity was verified byte-wise
(size identical, no non-PQ2_0 tensor touched).
metrics.jsonin this repo holds the raw aggregate counts.
ๆฅๆฌ่ชใกใข๏ผใใฌใใฅใผ็ใปใใฃใผใใใใฏๅ้๏ผ
prism-ml/Ternary-Bonsai-2-27B ใฎ PQ2_0 ใใใฏใใ้
ๅธใใฉใผใใใใฎใพใพ๏ผBF16 ใธ
ๆปใใใๅ้ๅญๅใใใ๏ผ2-bit ใณใผใใ็ดๆฅๆธใๆใใฆๆๅฆใ้คๅปใใใใฎใงใใใใกใคใซ
ใตใคใบใปใใณใฝใซๅฝข็ถใปใซใผใใซ็ต่ทฏใฏ่ฆชใใใฏใจๅฎๅ
จใซๅไธใงใ็ทจ้ๅฏพ่ฑกๅคใฎ 451 ใใณใฝใซใฏ
ใใคใๅไฝใงๅไธใงใใthinking ใฏ medium ใๆ็คบๆๅฎใใฆใใ ใใ๏ผๆขๅฎใฎ xhigh ใงใฏ
ๆๅฆ้คๅปใๅผฑใใๅ็ญใ็ฉบใซใชใๅ ดๅใใใใพใ๏ผใ
ใใใฏ v0.1 ใฎใใฌใใฅใผใงใใ ไธใฎ่กจใฏ็งใฎๅฐใใชๆค่จผใใผใในใฎ็ตๆใชใฎใงใ็ใใใฎ ็ฐๅขใปใใญใณใใใป่จ่ชใง่ฉฆใใ็ตๆใใใฒ Community ใฟใใฎ Discussion ใซๆธใใฆใใ ใใใ ็นใซใๆๅฆใๆปใไพใใ่ฝๅๅฃๅใใmulti-turn ใงใฎๅดฉใใใxhigh ใง็ฉบๅ็ญใซใชใไพใใๆญ่ฟ ใใพใใๆๆณใฎ่ฉณ็ดฐใฏใๆค่จผใใใๅฐใ้ฒใใ ๆฎต้ใงๅ ฑๆใใพใใ
- Downloads last month
- 24,065
2-bit