科利布里已问世:一款自主可控的开放权重模型
Kolibri Has Landed: A Sovereign Open-Weight Model

原始链接: https://aleph-alpha.com/en/blog/kolibri-has-landed-a-sovereign-open-weight-model/

Aleph Alpha 发布了 **Kolibri**,这是一款面向受监管企业和公共部门工作负载的自主可控英德双语混合专家模型。它拥有 **781 亿总参数**,每个词元约激活 **35 亿参数**,支持最长 **100 万词元**的上下文,并可调节推理强度。完整模型权重已在 Hugging Face 上依据 **Apache 2.0 许可证**发布。 Kolibri 的定位是作为更大型开放权重模型的低成本替代方案,在德语、数学、编程、长上下文、工具使用以及行业专用智能体任务方面表现出色。基于事实依据的训练使模型能够在文档无法支持答案时拒绝作答,从而减少幻觉。 一个自动化、基于代码的“模型工厂”支持开展数百项消融实验、频繁的检查点评测,以及从硬件和网络连接故障中自动恢复。训练使用了近 24 万亿个词元,其中包括大量自然生成的德语数据,以及一种感知词形的双语词元化器。Kolibri 根据欧洲和德国法律开发,具备完整的供应链透明度,并支持灵活的本部署署。

Hacker News 最新 | 往期 | 评论 | 提问 | 展示 | 职位 | 提交 登录 Kolibri 发布了:一款主权开放权重模型 (aleph-alpha.com) 25 分 由 bastitx 提交 3 小时前 隐藏 | 往期 | 收藏 | 2 条评论 帮助 mistyvales 0 分钟前 | 下一条 [–] Sega 32X 的游戏?? 回复 cyanydeez 37 分钟前 | 上一条 [–] 有意思的是,他们在推荐高端软件时没有考虑 4 位或 8 位量化,而且仍然使用 A3B;A3B 在廉价硬件上本应能提供不错的吞吐量。 如果他们能效仿 Qwen3.8-Flash-Next,或许可以大幅降低显存需求。 回复 可以考虑申请 YC 2027 年冬季批次! 申请截止至 11 月 2 日。 指南 | 常见问题 | 列表 | API | 安全 | 法律声明 | 申请 YC | 联系我们 搜索:
相关文章

原文

Research

Aleph Alpha

What Kolibri Delivers

Foundational capabilities for enterprise and government

Average benchmark score [%]
  • Kolibri
  • Kolibri Origin
  • Other post-trained models
  • Pareto frontier
Performance vs. throughput for post-trained models in English (left) and German (right). Metrics show unweighted average benchmark scores against decoded text per second and GPU. Higher and further right is better.
AIME 2025MathAIME 2025 (DE)MathAIME 2026MathAIME 2026 (DE)MathGPQA (diamond)KnowledgeGPQA (diamond, DE)KnowledgeAA-Omniscience IndexGrounding / hallucinationspublic set; from −100 to 100BrowseCompAgenticτ³-bench bankingAgenticτ²-bench retailAgenticτ²-bench airlineAgenticτ²-bench telecomAgenticBFCL v4 overallAgenticLiveCodeBench v6CodeHumanEval+CodeLongBench ProLong contextAA-LCRLong context
Show the numbers
BenchmarkKolibriKolibri OriginQwen3.6-35B-A3BNemotron 3 Super 120B-A12BMistral Small 4 119B-A6B
AIME 202596.981.984.691.779.8
AIME 2025 (DE)87.573.582.985.672.3
AIME 202696.081.591.090.483.1
AIME 2026 (DE)90.075.284.487.578.5
GPQA (diamond)84.368.183.478.074.7
GPQA (diamond, DE)81.358.580.676.672.9
AA-Omniscience Index-32.8-64.0-15.3-36.5-24.0
BrowseComp29.44.426.929.1–
τ³-bench banking38.15.710.615.55.7
τ²-bench retail69.958.571.667.562.9
τ²-bench airline76.758.770.772.740.0
τ²-bench telecom94.767.599.168.141.5
BFCL v4 overall61.436.467.261.058.0
LiveCodeBench v685.959.282.582.071.2
HumanEval+92.776.892.894.792.8
LongBench Pro64.5–70.862.956.4
AA-LCR68.3–69.767.052.3
Foundational capabilities of Kolibri. Kolibri is a balanced generalist model with competitive performance across math, code, long context, agentic capabilities, and knowledge. Benchmark scores are shown on a shared 0–100 scale, higher is better.

Contextualized performance for real-world applications

Score on internal customer-proxy benchmark
  • Automotive supplier 0.72 → 0.99
  • Semiconductors 0.35 → 0.80
  • German public sector 0.54 → 0.75
  • Industrial drive technology 0.31 → 0.60
  • Aerospace 0.14 → 0.59
  • one checkpoint, one eval
  • mean of that day
  • Kolibri Origin
  • Kolibri
Contextualized performance across successive post-training runs on internal customer-proxy benchmarks. Dots are single evaluations of training checkpoints, lines are the mean of each day's evaluations. Performance climbs across all five verticals, higher is better.

Control and compliance by design: upholding our customers' sovereignty

How We Built Kolibri at High Velocity

Our Model Factory: two models, three months apart

Kolibri Origin Kolibri
Finished pre-training11 June 202611 September 2026
Releaseno public release3 October 2026
Reasoning modeYes (one mode only)Yes (none, low, medium, high)
Total parameters30.6B78.1B
Active parameters / token3.27B3.46B
Pre-training tokens7.51T20T
Layers50 (2 dense + 48 MoE, 1 shared expert)50 (all MoE, 1 shared expert)
Pre-training context length8,192 (8k)16,384 (16k)
Longest trained length65,536 (64k)262,144 (256k)
Tokenizer vocabulary96,000128,000
Model dimension2,0482,560
Attention heads (query / KV)32 / 448 / 4
Experts (total / active)128 / 8384 / 6
Expert hidden dim768512
Attention patternfull attention, all layerssliding window (512) + full attention every 5th layer
Knowledge cutoffEN: 1 Sept 2024, DE: 1 Aug 2025EN/DE: 18 Jun 2026

Architecture and pre-training

Post-training at scale

German pre-training data

Specialized tokenizers for English-German

German web (FineWeb-2)

  • Kolibri 128,000 vocab 4.90
  • Kolibri Origin 96,000 vocab 4.69
  • Plain BPE 128k 128,000 vocab 4.89
  • GPT-5 200,019 vocab 4.35
  • DeepSeek V4 129,280 vocab 3.72
  • Kimi K3 163,586 vocab 3.28
  • GLM 5.3 154,856 vocab 3.93
  • Qwen3-Next 151,669 vocab 3.59
  • Qwen3.5-3.8 248,077 vocab 4.17
  • Gemini 262,144 vocab 4.13
  • EuroLLM 128,000 vocab 4.08
  • Tekken (Mistral, Nemotron, Apertus) 131,072 vocab 4.03

English web (FineWeb)

  • Kolibri 128,000 vocab 4.58
  • Kolibri Origin 96,000 vocab 4.50
  • Plain BPE 128k 128,000 vocab 4.59
  • GPT-5 200,019 vocab 4.67
  • DeepSeek V4 129,280 vocab 4.59
  • Kimi K3 163,586 vocab 4.62
  • GLM 5.3 154,856 vocab 4.61
  • Qwen3-Next 151,669 vocab 4.52
  • Qwen3.5-3.8 248,077 vocab 4.47
  • Gemini 262,144 vocab 4.49
  • EuroLLM 128,000 vocab 4.16
  • Tekken (Mistral, Nemotron, Apertus) 131,072 vocab 4.45
Tokenizer compression in average bytes per token on German and English web text. Kolibri achieves the best German compression in this comparison. Plain BPE 128k is standard BPE trained on the same data with the same settings as Kolibri, which isolates the effect of the training method. More text per token means fewer tokens per task – higher is better.

Bundessozialgerichtes Federal Social Court (genitive)

  • Kolibri Bundes sozial gericht es
  • GPT-5 Bund ess oz ial gericht es
  • Qwen3.8 Bund ess oz ial gericht es
  • Gemini Bund ess oz ial gericht es
  • Mistral Medium 3.5 · Nemotron 3 Nano Bund ess oz ial gericht es

silkworm

  • Kolibri silk worm
  • GPT-5 sil kw orm
  • Qwen3.8 sil kw orm
  • Gemini sil kw orm
  • Mistral Medium 3.5 · Nemotron 3 Nano sil kw orm

Protokolldaten log data

  • Kolibri Protokoll daten
  • GPT-5 Pro tok ol ld aten
  • Qwen3.8 Protokol ld aten
  • Gemini Protok ol ld aten
  • Mistral Medium 3.5 · Nemotron 3 Nano Pro tok ol ld aten

coprocessors

  • Kolibri co processors
  • GPT-5 cop rocess ors
  • Qwen3.8 cop rocess ors
  • Gemini cop rocess ors
  • Mistral Medium 3.5 · Nemotron 3 Nano cop rocess ors
How different tokenizers split the same words. The Kolibri tokenizer follows the morphology of the language. English and German words split into meaningful units, while competitors cut across morpheme boundaries. Lower is better: fewer, cleaner splits per word mean fewer tokens.

Grounding: reducing hallucinations

AA-Omniscience Non-Hallucination Ratepublic setshare not answered wrongRGB: holds backwhen the documents don'tanswer the questionRGB: invents nothingno falsehoods when thedocuments lack the answerFRAMESmulti-document reasoningM/A grounding scoreour own metricaxis from 0 to 0.5
Show the numbers
BenchmarkKolibriKolibri OriginQwen3.6-35B-A3BQwen3-Next 80B-A3BNemotron 3 Super 120B-A12BMistral Small 4 119B-A6B
AA-Omniscience Non-Hallucination Rate44.014.856.712.313.934.7
RGB: holds back85.673.979.681.374.682.3
RGB: invents nothing87.375.684.383.986.087.0
FRAMES71.265.774.768.974.971.9
M/A grounding score0.230.000.120.000.000.06

Contextualized performance

MuSiQue (cleaned)Agentic RAGHoneypotAgentic RAGSemiconductorsCustomer proxyGerman public sectorCustomer proxyAerospaceCustomer proxyAutomotive supplierCustomer proxyIndustrial drive technologyCustomer proxy
Show the numbers
BenchmarkKolibriKolibri OriginQwen3-Next 80B-A3BQwen3.6-35B-A3BNemotron 3 Super 120B-A12BMistral Small 4 119B-A6B
MuSiQue (cleaned)77.342.750.561.279.166.8
Honeypot80.825.313.574.368.868.1
Semiconductors80.435.341.279.469.662.7
German public sector75.054.029.572.078.050.0
Aerospace58.914.148.159.054.947.0
Automotive supplier99.072.484.292.691.087.1
Industrial drive technology60.031.432.759.537.356.8

Benchmarks

Type MoE MoE MoE Dense
Active parameters 3B 4–6B 12B 27B · 70B
Benchmark Kolibri Kolibri Origin GLM-4.7 Flash 30B-A3B Nemotron 3 Nano 30B-A3B Qwen3.5 35B-A3B Qwen3.6 35B-A3B Qwen3-Next 80B-A3B Thinking Gemma 4 26B-A4B IT GPT-OSS 120B Mistral Small 4 119B-A6B GLM-4.5 Air 106B-A12B Nemotron 3 Super 120B-A12B Qwen3.8 27B Apertus 70B Instruct
Overall (EN) 75.5 54.1 64.7 65.6 74.7 71.4 62.4 71.9 72.3 63.1 64.4 73.0 80.2 –
Overall (DE) 70.8 46.4 50.4 59.3 69.8 67.3 58.0 66.3 70.2 61.4 64.8 67.9 79.9 –
Knowledge
Average (EN) 50.1 39.7 45.5 46.0 52.7 52.1 48.4 51.4 50.0 47.5 45.7 52.0 56.8 –
Average (DE) 57.6 44.9 46.5 41.2 61.3 61.0 55.7 61.5 58.0 51.4 52.2 59.5 69.2 –
GPQA Diamond (EN) 84.3 68.1 73.1 73.9 83.8 83.4 76.1 81.1 76.4 74.7 73.2 78.0 89.2 29.5
GPQA Diamond (DE) 81.3 58.5 59.8 49.6 84.2 80.6 72.2 80.1 76.0 72.9 71.1 76.6 88.1 31.4
Humanity's Last Exam (EN) 21.5 9.4 15.4 12.1 20.4 21.1 11.6 19.2 19.4 9.7 8.7 20.6 35.6 5.2
Humanity's Last Exam (DE) 15.9 10.4 9.1 13.1 18.1 20.5 15.6 23.4 20.7 10.5 10.5 22.3 37.2 5.7
AA-Omniscience Accuracy (public set) 14.8 11.3 17.0 19.5 22.0 19.5 24.2 20.7 23.3 25.0 20.0 26.7 17.5 13.5
AA-Omniscience Index (public set) −32.8 −64.0 −62.8 −45.7 −47.3 −15.3 −42.3 −47.3 −35.2 −24.0 −28.8 −36.5 −9.5 –
MMLU-Pro CoT (EN) 80.0 70.1 76.5 78.3 84.6 84.3 81.7 84.5 80.8 80.4 80.9 82.7 85.0 43.0
MMLU-ProX CoT (DE) 75.5 65.7 70.7 61.0 81.7 81.9 79.4 81.1 77.2 70.7 74.9 79.7 82.4 37.3
Math
Average (EN) 96.5 81.7 88.8 88.8 90.1 87.8 86.3 87.4 90.7 81.4 82.8 91.1 97.8 –
Average (DE) 88.8 74.3 45.2 84.3 79.6 83.7 83.7 88.1 90.8 75.4 80.6 86.5 96.7 –
AIME 2025 (EN) 96.9 81.9 89.4 89.6 88.1 84.6 84.2 87.3 90.6 79.8 81.9 91.7 97.9 0.6
AIME 2025 (DE) 87.5 73.5 43.8 84.4 76.7 82.9 80.6 88.1 90.6 72.3 80.6 85.6 96.5 0.2
AIME 2026 (EN) 96.0 81.5 88.3 87.9 92.1 91.0 88.5 87.5 90.8 83.1 83.8 90.4 97.7 0.6
AIME 2026 (DE) 90.0 75.2 46.7 84.2 82.5 84.4 86.7 88.1 91.0 78.5 80.6 87.5 96.9 0.0
Agentic
Average (EN) 63.4 41.6 58.9 46.4 63.4 62.1 46.3 54.6 54.0 40.7 53.5 54.9 66.7 –
TerminalBench 2.1 27.7 – 20.2 9.7 39.7 – 8.6 – 29.2 21.0 – 39.7 76.8 –
Tau2-Bench (Telecom) 94.7 67.5 95.9 45.9 97.7 99.1 43.9 45.3 73.1 41.5 53.8 68.1 82.5 10.8
Tau2-Bench (Retail) 69.9 58.5 57.9 64.9 70.8 71.6 60.8 71.3 60.5 62.9 61.4 67.5 68.7 9.6
Tau2-Bench (Airline) 76.7 58.7 68.7 52.7 76.0 70.7 65.3 73.3 72.7 40.0 70.7 72.7 83.3 40.0
Tau3-Bench (Banking) 38.1 5.7 7.2 5.7 11.3 10.6 5.4 16.0 14.7 5.7 6.4 15.5 50.0 2.1
BFCL v3 (multi-turn) 39.8 22.8 58.2 47.9 54.0 53.5 51.4 53.4 45.6 36.2 61.6 44.6 42.5 0.6
BFCL v4 (overall) 61.4 36.4 65.4 61.5 70.5 67.2 51.0 68.2 57.3 58.0 67.2 61.0 73.2 –
BFCL v4 (non-live AST) 79.1 78.1 83.3 85.0 85.8 88.2 83.6 83.7 35.8 83.6 85.5 45.0 85.3 –
BFCL v4 (live) 78.9 73.7 78.3 78.8 80.2 81.4 82.5 80.2 70.4 78.4 78.2 77.6 79.9 –
BFCL v4 (multi-turn) 47.5 27.5 62.7 53.5 59.9 58.1 56.0 61.4 55.4 40.4 65.2 51.7 55.5 –
BFCL v4 (memory) 62.8 19.4 41.5 39.1 62.6 53.8 35.3 52.9 50.7 39.1 43.4 59.6 79.6 –
BFCL v4 (web search) 62.5 10.5 69.0 66.0 75.0 68.5 12.5 75.0 57.0 69.0 71.0 71.5 82.0 –
BrowseComp 29.4 4.4 – 14.5 36.5 26.9 2.8 25.5 31.2 – – 29.1 46.4 –
Code
Average (EN) 89.3 68.0 67.8 81.8 85.0 87.7 83.6 89.0 90.8 82.0 79.8 88.3 94.2 –
LiveCodeBench v6 85.9 59.2 46.5 71.3 77.8 82.5 73.9 82.3 87.5 71.2 67.8 82.0 93.8 8.7
HumanEval+ 92.7 76.8 89.0 92.4 92.2 92.8 93.3 95.7 94.1 92.8 91.8 94.7 94.7 41.6
SWE-Bench Verified 66.4 – 51.0 38.6 71.6 73.8 – 57.8 – 60.8 11.6 60.2 72.6 –
Instruction Following
Average (EN) 78.1 62.5 64.5 73.2 72.7 66.1 60.7 79.9 71.1 49.8 38.2 73.7 81.9 –
IFBench (loose-prompt) 78.1 62.5 64.5 73.2 72.7 66.1 60.7 79.9 71.1 49.8 38.2 73.7 81.9 25.5

Get Started

pip install "aleph-alpha-inference>=1.0"
vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \
  --reasoning-parser kolibri1 \
  --tool-call-parser kolibri1 \
  --enable-auto-tool-choice

Contact for Deployment and Specialization

联系我们 contact @ memedata.com