运行 1.58 位(BitNet)语言模型的 ESP32S3 集群
ESP32S3 cluster running 1.58-bit (BitNet) Language model

原始链接: https://github.com/Low-Zi-Hong/ESP32s3-LLM-Cluster

本项目实现了一个分布式推理引擎,在七个 ESP32S3 微控制器组成的集群上运行 0.5B 参数、1.58 位 (BitNet) 的语言模型。 该架构采用一个主节点负责分词、嵌入和最终采样,其余六个节点负责处理 Transformer 模块。计算在节点间按顺序分布,并通过高速菊花链 SPI 总线进行通信。每个计算节点都针对 BitNet 推理进行了优化,具备用于 1.58 位三值线性层的汇编级 MAC(乘累加)运算功能,以及 PSRAM 中的 KV 缓存管理。 项目仓库包含: * **主节点固件:** 处理调度、用户输入输出以及嵌入查找。 * **节点固件:** 用于并行执行 Transformer 层的优化工作节点。 * **Python 工具链:** 一套完整的模型量化 (QAT)、词汇表剪枝、权重打包和二进制序列化套件。 该系统专为资源受限的硬件设计,展示了极低位量化和分布式并行处理如何使小型嵌入式集群实现大模型推理。有关接线、烧录和模型准备的详细说明,请参阅项目中的 `workflow.md`。

Hacker News 最新 | 过往 | 评论 | 提问 | 展示 | 招聘 | 提交 登录 在 ESP32S3 集群上运行 1.58 位 (BitNet) 语言模型 ( github.com/low-zi-hong ) 36 点 由 nkko 发布 5 小时前 | 隐藏 | 过往 | 收藏 | 3 条评论 帮助 tdhz77 11 分钟前 | 下一条 [–] 很快每个灯泡里都会运行 Kubernetes 的人工智能了 回复 cameron_b 1 小时前 | 上一条 | 下一条 [–] 看到这种“压缩”程度使其变成了一个华而不实的语言模型噪声发生器,确实有点扫兴。不过它依然很迷人。 回复 sjakati98 59 分钟前 | 上一条 [–] Gemma 4 什么时候出? 回复 指南 | 常见问题 | 列表 | API | 安全 | 法律 | 申请 YC | 联系 搜索:
相关文章

原文

A distributed pipeline inference engine on multiple ESP32S3 running 1.58-bit (BitNet) Language model.

ESP32S3 boards

This project runs a sliced 0.5B LLM across a cluster of 7 ESP32s3. One act as master and others are node. The master node runs the tokenizer and embeding and the other attention layer and MLP ran on the nodes. The master and node communicate through high speed SPI Daisy-Chain.

┌─────────────────────────────────────────────────────────┐
│                     MASTER NODE                         │
│                                                         │
│  [ Prompt ] ---> BPE Tokenizer                          │
│                       │                                 │
│                 Token Embedding                         │
│             (INT4, ~14MB in Flash)                      │
│                       │                                 │
│             (SPI CH A - TX to Node 1)                   │
└───────────────────────┬─────────────────────────────────┘
                        │ Hidden State Vector (FP32)
                        ▼
┌─────────────────────────────────────────────────────────┐
│                    COMPUTE NODE 1                       │
│             (SPI CH B - RX from Master)                 │
│                                                         │
│  ► Layer 0 to 3 (4x Transformer Blocks)                 │
│    • RMSNorm (FP16 scaled to FP32)                      │
│    • 1.58-bit Attention (Q, K, V, O proj) + RoPE        │
│    • KV Cache (PSRAM)                                   │
│    • 1.58-bit MLP (Gate, Up, Down proj)                 │
│                                                         │
│             (SPI CH A - TX to Node 2)                   │
└───────────────────────┬─────────────────────────────────┘
                        │
                       ... (Nodes 2 to 5)
                        │
                        ▼
┌─────────────────────────────────────────────────────────┐
│                    COMPUTE NODE 6                       │
│             (SPI CH B - RX from Node 5)                 │
│                                                         │
│  ► Layer 20 to 23 (4x Transformer Blocks)               │
│    • Same 1.58-bit Architecture                         │
│                                                         │
│             (SPI CH A - TX back to Master)              │
└───────────────────────┬─────────────────────────────────┘
                        │
                        ▼
┌─────────────────────────────────────────────────────────┐
│                     MASTER NODE                         │
│             (SPI CH B - RX from Node 6)                 │
│                                                         │
│                 Final RMS Norm                          │
│             (FP16, 64KB in 'fnorm' partition)           │
│                       │                                 │
│         LM Head (Tied to INT4 Embeddings)               │
│                       │                                 │
│               Greedy Sampling                           │
│                       │                                 │
│  [ Output ] <--- Next Token ID                          │
└─────────────────────────────────────────────────────────┘

pls refer workflow guide to start with the project.

.
├── README.md                   # Project documentation
├── workflow.md                 # Step-by-step flashing, model prep & wiring guide
├── .gitignore                  # Git ignore rules for build files & binaries
│
├── docs/                       
│   └── images/                 # Architecture diagrams and hardware photos
│
├── master_board/               # Firmware for the Master Node (ESP-IDF)
│   ├── main/
│   │   ├── main.cpp            # Master orchestrator, user I/O & BPE tokenizer
│   │   ├── embedding.cpp       # INT4 embedding lookup logic
│   │   ├── lm_head.cpp         # LM Head mapping and greedy sampling
│   │   └── spi_bus.cpp         # Master dual-channel SPI driver
│   ├── partitions.csv          # Custom partition table (token, model, fnorm)
│   └── CMakeLists.txt
│
├── node_firmware/              # Firmware for the Compute Nodes (ESP-IDF)
│   ├── main/
│   │   ├── main.cpp            # Node worker entry point & inference loop
│   │   ├── bitlinear.cpp       # 1.58-bit ternary linear layer implementation
│   │   ├── bitlinear_forward.S # Assembly optimized MAC ops for 1.58-bit
│   │   ├── qwen_attention.cpp  # Qwen Attention, RoPE & KV-Cache runtime
│   │   ├── lut_table.cpp       # Look-up tables for extreme optimization
│   │   └── spi_bus.cpp         # Daisy-chain SPI DMA receiver/transmitter
│   ├── partitions.csv          # Layer partition layout for Node
│   └── CMakeLists.txt
│
├── python_tools/               # PC-side quantization & preprocessing suite
    ├── crop_token.py           # Vocabulary pruning (scales down to 32K tokens)
    ├── crop_model_weight.py    # Embedding matrix slicing
    ├── qat_158.py              # BitNet QAT (Quantization-Aware Training) fine-tuning
    ├── bit4_embedding.py       # INT4 weight packer for embeddings
    ├── pack_tokenizer_bin.py   # Serializes tokenizer rules into ESP32 .bin
    ├── pack_model_bin.py       # Packs 1.58-bit layer chunks for physical alignment
    ├── look_model_structure.py # Debug tool for inspecting .safetensors
    └── flash_*.bat             # Multi-threaded fast flashing scripts

This project is licensed under the MIT License - see the LICENSE file for details.

Inspiration, related works, and references:

联系我们 contact @ memedata.com