Models

Conway Research Releases Saluki 27B for Local AI Agents

Conway Research has released Saluki 27B, a compressed open-source model that outperforms its 54 GB base model at local tool calling while running in just 7.89 GB of memory.

MarkTechPost2 days agoModels
Image: MarkTechPost

Conway Research's Underdog team has released Saluki 27B under an Apache 2.0 license, giving developers a compact agent model that outperforms its uncompressed original at tool calling. Built as a 2-bit mixed-precision GGUF quantization of Alibaba's 27-billion parameter Qwen3.8-27B model, Saluki 27B drops the required file size from 54 GB in BF16 down to 7.89 GB. The model runs on stock llama.cpp with full GPU offloading and supports optional vision add-on modules at 629 MB or 928 MB.

To achieve this size reduction while preserving function-calling abilities, Underdog stacked three layers of development. The base Qwen3.8-27B model features 64 layers combining Gated DeltaNet linear attention with gated attention and natively supports 262,144 tokens. Next, ISTA-DASLab applied Grouped Scalar Quantization and Rate-Constrained Optimization to create Qwen3.8-27B-GSQ-RCO-GGUF, producing an 8.4 GB IQ2_XS build at 2.50 bits per weight. Underdog then executed a third custom pass, producing an IQ2-mix file with an imatrix tag that shrank the footprint to 7.89 GB while specifically targeting tool retention.

Across nine benchmarks, Saluki 27B retains an average of 96% of the original model's performance while scoring higher on key agent tasks. On the Underdog Bench—comprising 120 frozen tasks from BFCL v4 evaluated with thinking disabled at temperature zero—Saluki scored 88 compared to 84 for full Qwen3.8-27B and 70 for PrismML's 5.95 GB Bonsai 2. In 100 parallel tool calling tests, Saluki recorded 42 successes against 35 for the full model. On SWE-bench Verified, Saluki resolved 30 out of 50 software issues compared to 33 for the baseline.

However, heavy quantization reduced capacity in mathematics and multi-step reasoning. Saluki logged 79.2 on AIME 2025 (avg@4) versus 96.7 for the original, and 80.0 on AIME 2026 versus 94.6. On MBPP+ it registered 78.0 against 83.9, and on MuSR it scored 67.5 against 79.6. Prompt-loose scores showed slight gains on IFEval (93.5 vs 91.5) and IFBench (72.7 vs 71.0). Unlike competing ultra-low-bit models like PrismML Bonsai 2 27B—which achieves 1.75 bits per weight but requires a custom engine fork—Saluki executes directly in standard open-source runtimes.

This is our own summary of reporting by MarkTechPost

More in Models