Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up

All HF Hub posts

danielhanchenΒ 
posted an update 2 days ago
view post
Post
5348
Unsloth has surpassed 500M model downloads on Hugging Face! πŸ¦₯πŸ€—

Qwen3.8-27B GGUF is already Unsloth’s #1 most-downloaded model ever.

Thanks for all your support! Follow us:
unsloth
Banaxi-TechΒ 
posted an update 2 days ago
view post
Post
4497
We're excited to release BananaAll, our SLM Super App.

It allows you to do EVERYTHING you need to do to trains SLMs in a single app, no terminal, no 30 chrome tabs.

The train tab allows you to train models, select datasets from presets, and use other ones with auto mapping, model size slider, it automatically generates a training script for you.

Then after you've trained the model or want to compare it to competitors, the evaluation tab, run ARC EASY, ARC Challenge, Hellaswag, PIQA, Arithmark 3, BananaMind Base Bench and more! Simple Results screen.

And lastly the inference tab, run your trained models or others.

Normally you would need seperate apps or scripts for that, but the BananaAll Super App lets you do all of that in a single app.

We also trained a small 2.5M parameter model on 200M tokens of Fineweb edu, The results: BananaMind Base Bench 854 and 53% on PIQA. On only 200M tokens.

Check it out at https://github.com/BananaMind/BananaAll.
  • 23 replies
Β·
SeaWolf-AIΒ 
posted an update 2 days ago
view post
Post
3575
πŸ–ΌοΈ NO GPU, Only CPU : Z-Image model

Zero graphics cards. 46 seconds. Photoreal.

That laptop you're reading this on. No graphics card, right? It generates images.

No CUDA install. No Python environment. No driver changes. One binary, three model files. Done.

πŸ“Š Measured β€” GPU count used: zero

512Γ—512 : 46.4 s
Korean prompt : 45.3 s (faster than English)
1024Γ—1024 : 192.7 s
Peak RAM : 6.42 GB
GPUs used : 0

(Intel Xeon Gold 6526Y Γ—2, 48 threads, Q4_0, 3 steps)

⚑ From 244 seconds to 46 β€” 5.3Γ—

Run it on defaults and it takes 244 s. Switch to 3 steps and it's 48.6 s. Add VAE tiling and it's 46.4 s.

The biggest culprit was the default. Z-Image Turbo is distilled to paint in few strokes, but the tool's default is 20. We were throwing away 5Γ— for no reason. So were we, at first.

3 is the floor. Put 4 and 3 side by side and you cannot tell them apart. At 2 it collapses β€” water droplets and wood grain vanish, and the surface turns cloth-like.

πŸ”— Links

Model
FINAL-Bench/POCKET-Zimage-CPU

Live demo Space (runs on CPU)
FINAL-Bench/POCKET-Zimage-CPU

POCKET collection
FINAL-Bench/pocket-models-6a618ee5d23eafb7e185a5c6
  • 2 replies
Β·
erenikoΒ 
posted an update 5 days ago
view post
Post
227
NEW CLAUDE HAIKU MODEL COMING

FINALLY AFTER A YEAR

ANTHROPIC ANNOUNCED HAIKU 5.5
EnderchefΒ 
posted an update 3 days ago
view post
Post
2911
AxiomicLabs released new benchmark, Tiny Theory of Mind, to test your SLM models' Theory of Mind Intuition!
Check it out and like it!

AxiomicLabs/Tiny_Theory_of_Mind
Banaxi-TechΒ 
posted an update about 2 hours ago
view post
Post
346
We're releasing a MAJOR update to the BananaAll SLM Super App.
If you want to use a custom architecture, previously you had to go trough reviewing the code yourself, now add an Openrouter API key and review it with GPT 6 Luna in one button. A review cost be half a cent so anyone can try it. This is one of the main features.
Now ROCm, AMD and Windows, Mac support.
Colab and Molab support.

Detailed list of features:
Get improved Windows Python detection and support paths for compatible AMD ROCm, Intel XPU, and Apple MPS setups.
Choose local training or export a self-contained Python script for Colab or Molab. Notebook runs produce a downloadable model ZIP.
Start pretraining with an existing model’s tokenizer, or train a new one from your datasets.
Try experimental 1.58-bit Ternary fake-quantized training on NVIDIA GPUs.
Watch live tokens per second. Model compilation is on by default and falls back automatically if it fails.
Build custom architectures with separate configuration and modeling files, then review the training code manually or with optional OpenRouter AI Review.
Install from source with the new coding-agent instructions.
This release also fixes inflated loss reporting for custom models.



And for those users who didn't want to try it out just because installation would be so hard, it isnt now.
Go to any coding agent (Pi, Claude Code, Codex, OpenCode, basically all work), and just paste "Install BananaAll for me. Fetch and follow https://raw.githubusercontent.com/BananaMind/BananaAll/main/agent_install.txt."
That's it.

Check it out at https://github.com/BananaMind/BananaAll/

Also on SAICR, we're currently training a new major model (NACR v2) and ACR 1.0 is in the finishing.

  • 2 replies
Β·
Hoglet-33Β 
posted an update 3 days ago
view post
Post
2269
Everything going on here at basically AI:

1. Pebble 1.5

We're working on Pebble 1.5. Here's what we know so far:

- They will be better than the last generation. 99.99% certain.
- Expanded context lengths of at least 16,384 tokens, with the flagship potentially reaching 32,768.
- A Mamba3-based architecture with some other new architectural designs we're experimenting with.
- Native CPU compatibility β€” something we failed at with the last generation.
- Natively multilingual and multimodal???

2. SmolCodeBench

A code benchmark designed specifically for small models, because there really isn't a good one right now.

3. SENTRY

VOID is working on something called SENTRY β€” System for Evaluating Neural Threats, Responses, and Yields.

More on that soon.

4. basically OS

It's an operating system/app/harness. We're still deciding.

5. Finances

Trying to balance the finances after purchasing a Hugging Face Pro subscription.

Follow us for updates:
@Hoglet-33
basically-ai

basically-experimental

void-research
  • 2 replies
Β·
sharpenbΒ 
posted an update 3 days ago
view post
Post
1752
Today, we open-source Pruna-Qwen-Image-2.1, a set of a few-step LoRA adapters that make Qwen-Image-2. up to 6.3Γ— faster for image generation and editing.

Try it here: PrunaAI/Pruna-Qwen-Image-2.1
DavidAUΒ 
posted an update 17 days ago
view post
Post
14529
Qwen 3.8 27B - TWIN TURBO, Fable Fusion (10 modes of operation)

Tuned, and tweaked to match the legendary Qwen 3.6 27B FF711 (2300+ likes, 4 million+ downloads) this fine tune matches the stability and power at "arc-c" 709: (118 pts higher than Qwen 3.8 27B) (The OpenAI, Claude and Gemini "zone of intelligence") in 8 bit and 701 arc-c in 4 bit AND THIS is instruct mode - thinking/reasoning is higher.

This version is called TWIN-TURBO because it drastically reduces thinking tokens (by 1/2 to as LOW as 1/20), yet maintains output detail and quality. In other words while "reg" Qwen3.8 27B is thinking about "formatting" for a few 1000 tokens, this model is already done and waiting for more.

This repo contains both "regular" and "MTP" Neo and NEO MAX GGUF quants.

BUT WE WENT FURTHER:

Now with 5 reasoning modes (2 new - UltraXhigh / Einstein), and 5 instruct modes (2 new - UltraXhigh / Einstein, all use ZERO REASONING TOKENS) all switchable on the fly via API, direct and "in chat" (yes - model control at the chat/message level).

GGUFS:

DavidAU/Qwen3.8-27B-TWIN-TURBO-Fable-Cold-Fusion-709-L-Uncensored-NM-DAU-NEO-MTP-GGUF

SOURCE:

DavidAU/Qwen3.8-27B-TWIN-TURBO-Fable-Cold-Fusion-709-L-Uncensored

UPDATE:
This model (and a few more) will soon have 12 reasoning and 12 instruct modes plus interactive help, model embedded system to select the best reasoning/instruct mode[s] for all use cases.

Final testing is in progress...
  • 2 replies
Β·
harshitkguptaΒ 
posted an update about 11 hours ago
view post
Post
2533
Fine-tuned Qwen 2.5 (0.5B β†’ 3B) on real coding-agent traces, 10 controlled runs, one 16GB Mac. Compared PyTorch MPS vs. Apple MLX for local LoRA SFT β€” and the honest answer is "it depends on what you're optimizing for":

β€’ PyTorch MPS: 2.2x–5.7x faster raw throughput, but hits a hard memory wall β€” can't load a 3B model in FP16 on 16GB.
‒ Apple MLX: 4-bit QLoRA fits 3B+ models with almost flat memory scaling as context grows (+109 MB going from 1k→4k tokens).
β€’ 4-bit quantization doesn't cost you convergence β€” eval loss tracks closely across backends.
β€’ The bigger surprise: most of MLX's slowdown isn't the 4-bit dequant tax. Two of the 10 runs went unquantized to isolate it β€” dequant only explains 1.07x–1.4x of the gap. A ~4.1–4.6x framework-level gap remains either way.

All 10 LoRA adapters + Trackio logs are public so the numbers are checkable, not just claimed.

Full writeup: https://huggingface.co/blog/harshitkgupta/fine-tuning-coding-agents-on-mac-pytorch-mps-mlx