Ling-3.0-tiny โ€” MLX 4-bit

Run with Rapid-MLX: rapid-mlx serve ling-3.0-tiny-4bit starts an OpenAI- and Anthropic-compatible API on your Mac. ยท Rapid-MLX benchmark vs Apple's MLX

The first MLX conversion of inclusionAI/Ling-3.0-tiny: a 7.9B-total / 1.3B-active sparse-MoE reasoner (128 experts, top-8 + 1 shared) with a KDA + MLA hybrid attention stack and 131K context, MIT licensed.

4.2 GB at 4.507 bits/weight โ€” it fits and runs on an 8 GB Apple Silicon Mac.

Quantization 4-bit, group size 64 (router kept 8-bit, short-conv weights fp)
Size on disk 4.2 GB
Context 131,072 tokens
Active parameters 1.3B per token
License MIT (inherited from the base model)

Serve it

The bailing_hybrid architecture is not in upstream mlx-lm yet โ€” this checkpoint is served by rapid-mlx, which ships a verified native implementation (reference parity 1.5e-6 against the official modeling code):

pip install -U rapid-mlx   # 0.12.10 or newer
rapid-mlx serve ling-3.0-tiny-4bit

You get an OpenAI-compatible server on localhost:8000 with reasoning (reasoning_content) and tool calling parsed natively โ€” thinking is controlled with chat_template_kwargs: {"enable_thinking": true} or the model's detailed thinking on/off system-prompt switch.

Once mlx-lm gains native bailing_hybrid support, this checkpoint will load there unchanged.

Conversion provenance

Converted with mlx_lm.convert (quantize=True, q_bits=4, q_group_size=64) running rapid-mlx's vendored bailing_hybrid implementation (PR #1817), which was verified against the official modeling_bailing_moe_v3.py on identical random weights to a max logits deviation of 1.5e-6 (full prefill) / 1.9e-6 (token-by-token incremental) before conversion. End-to-end chat / reasoning / tool-call behaviour validated on an M2 Pro Mac mini.

Downloads last month
84,289
Safetensors
Model size
8B params
Tensor type
U32
ยท
BF16
ยท
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for rapid-mlx/Ling-3.0-tiny-MLX-4bit

Quantized
(32)
this model

Space using rapid-mlx/Ling-3.0-tiny-MLX-4bit 1