Show HN: InstinctFlash – Run 5B world-action models in real time on Jetson Thor

原始链接: https://github.com/General-Instinct/InstinctFlash

**InstinctFlash** 是一款用于机器人模型的统一运行时与加速框架,支持 NVIDIA Jetson Thor、RTX 4090 及 RTX 5090 平台。它通过单一、一致的 API 实现高性能推理,在 Jetson Thor 上利用 FP8 精度和优化采样技术,最高可带来 33.78 倍的速度提升。 **核心功能:** * **广泛兼容性:** 支持包括 LingBot-VA、pi0.5、GR00T 和 Cosmos3 在内的八个模型系列,并可轻松部署自定义微调检查点(checkpoints)。 * **优化服务:** 提供用于本地执行的 Python API,以及兼容现有机器人生态系统(如 openpi-client)的高性能 WebSocket 接口。 * **循证优化:** 该运行时采用六层优化策略,涵盖从权重压缩到特定硬件内核融合,确保性能提升可证明、可验证。 * **工具支持:** 包含用于模型引导、验证和基准测试的 CLI 工具。该框架支持用户验证所发布检查点的非劣效性,并通过 Rerun 实时传输遥测数据。 InstinctFlash 简化了从训练到部署的流程,开发者只需将运行时指向任意检查点,系统即可自动识别需求、生成执行计划并以最少的配置完成模型服务。

General Instinct 发布了 **InstinctFlash**,这是一个开源(AGPL-3.0)的高性能服务框架,旨在加速机器人模型。该框架旨在解决部署大型视觉-语言-动作(VLA)和世界动作模型时固有的延迟问题,可在 NVIDIA Jetson Thor、RTX 4090 和 5090 等硬件上实现实时推理。 InstinctFlash 利用六种主要的优化技术——包括 CUDA 图捕获、内核融合、内存规划和少步蒸馏——实现了显著的加速。在测试中,该框架在保持高成功率(90.5% 对比 92.1% 的基准)的同时,使 LingBot-VA 模型的性能提升了高达 33.78 倍。 该框架目前支持八个模型系列,包括 pi0.5 和 NVIDIA Cosmos Policy。用户可以通过 Python 运行时或兼容 OpenPI 的 WebSocket 服务器部署微调后的检查点。InstinctFlash 已被三星和西门子等公司的团队采用,现已在 GitHub 上向寻求优化控制回路的机器人开发人员公开。
相关文章

原文

  • [2026/09/17] RTX 5090 support. Deploy on your workstation with the same Runtime API used on Jetson Thor. Setup · Reproduce.
  • [2026/09/16] RTX 4090 support. Desktop inference and WebSocket serving with dedicated installation profiles. Setup · Reproduce.
  • [2026/09/15] Full-source release. Eight robotics model families, acceleration kernels, and Python / WebSocket serving through one Runtime. Get started.
  • [2026/09/15] Jetson Thor benchmarks. Up to 33.78× speedup with LingBot-VA @2V/4A, using FP8 and fewer sampling steps. Results · Reproduce.

Prediction p50 on Jetson Thor (ms), measured September 15, 2026.

We’ve seen up to 33.78× speedup with no observed loss in task performance in our real-robot tests.

Model Acceleration line PyTorch InstinctFlash Speedup
LingBot-VA FP8 · 25V/50A 15506.32 2891.74 5.36×
↳ LingBot-VA FP8 · 2V/4A 2071.29 459.10 4.51×
LingBot-VLA-4B FP8 624.22 221.53 2.82×
LingBot-VLA-V2-6B FP8 734.56 394.11 1.86×
Cosmos3 Edge DROID NUMERIC · UniPC4 / CFG3 3393.78 1048.01 3.24×
Cosmos3 Nano DROID NUMERIC · UniPC4 / CFG3 10184.68 4772.38 2.13×
pi05 FP8 408.58 51.85 7.88×
GR00T N1.7 BITEXACT 139.50 117.30 1.19×
DreamZero DROID FP8 · 16 steps · dynamic cache 23563.08 11899.42 1.98×

VA measures early continuations; each row compares the same schedule. The 33.78× headline includes 25V/50A → 2V/4A. FP8 and sampling changes are optional.

Protocol and raw results · Native VA 2V/4A · Reproduction commands

git clone https://github.com/General-Instinct/InstinctFlash && cd InstinctFlash
python3 -m venv .venv-core
source .venv-core/bin/activate
python -m pip install . uv==0.12.5

The Python 3.10+ core inspects checkpoints and plans without PyTorch or a GPU. Inference uses a separate, pinned environment for each model family. For RTX 4090:

python3 scripts/bootstrap_vendor.py install pi05 --target rtx4090 \
  --python python3.12 --root ~/ifl-pi05-4090 --ptxas /usr/local/cuda/bin/ptxas
source ~/ifl-pi05-4090/activate.sh

Use va, vla4, vla2, pi05, groot, edge, nano or dreamzero. Edge and Nano use Python 3.13; the other families use Python 3.12. The bootstrap installs the upstream source, compatibility patches, core and adapter. Model weights are downloaded separately. See RTX 5090 setup, RTX 4090 setup or Jetson Thor setup, which selects --target jetson_thor and uses the Thor CUDA backend build.

Your fine-tuned checkpoint — the expected case. Point serve at the training output; it detects the family, writes the small instinctflash.json declaration from what the checkpoint itself proves, and starts serving. One command:

instinctflash serve /path/to/your/checkpoint

Anything the checkpoint cannot prove is asked for explicitly, never guessed. Once the declaration exists (serve writes it on first run), the same directory also loads in Python:

from instinctflash import Runtime

runtime = Runtime.from_pretrained("/path/to/your/checkpoint")

A stock release — use its Hub id after installing the family's environment:

runtime = Runtime.from_pretrained("robbyant/lingbot-va-posttrain-robotwin")
family model id
LingBot-VA (5B WAM) robbyant/lingbot-va-posttrain-robotwin
LingBot-VLA-4B robbyant/lingbot-vla-4b-posttrain-robotwin
LingBot-VLA-V2-6B robbyant/lingbot-vla-v2-6b-robotwin
pi0.5 lerobot/pi05_base · lerobot/pi05_libero_finetuned_v044
GR00T-N1.7-3B nvidia/GR00T-N1.7-3B
Cosmos3 policies nvidia/Cosmos3-Edge-Policy-DROID · nvidia/Cosmos3-Nano-Policy-DROID
DreamZero GEAR-Dreams/DreamZero-DROID

Fine-tunes reuse their family's adapter; quality is evaluated per checkpoint.

The same Runtime defaults to precision="native" with a BITEXACT transformation ceiling. Use tier_ceiling="numeric" to allow numerical changes, or precision="fp8" (CLI: --fp8) to explicitly enable FP8. Step schedules are selected separately. See precision policy and FP8 support and validation.

DreamZero's opt-in dynamic step cache requires tier_ceiling="behavioral" with either precision. See the Thor measurements.

In process — this is the whole Python API:

with runtime.episode(prompt="put the bottle in the dustbin") as episode:
    while not done:
        result = episode.predict(observation)
        action = result["action"]

observation is a dict in the model's own format; result["action"] contains its action array. For LingBot-VA, pass executed_action=... when the controller changes a predicted action chunk, so the next prediction uses the actions actually executed.

Over the network — the serve command above hosts the same runtime behind the msgpack-over-websocket wire protocol the pi0/openpi ecosystem already speaks, so existing robot-side clients connect unchanged (pip install openpi-client):

from openpi_client.websocket_client_policy import WebsocketClientPolicy

client = WebsocketClientPolicy("my-server", 8000)
result = client.infer(observation)
action = result["action"]

The prompt rides in the observation; a changed prompt starts a new episode, and a client can say it explicitly with {"reset": True, ...}. Four flags cover the rest:

  • --serve.dry_run — preflight only: device, declaration, plan. No weights, no GPU.
  • --serve.smoke — load, produce one action, exit.
  • --serve.seed — seed native execution for paired comparisons; FP8 serving rejects this option.
  • --serve.viz — stream observations, actions and latency to a Rerun viewer.

The second verb, instinctflash validate <dir>, checks a checkpoint is publishable; given --validate.teacher_outcomes/.student_outcomes/.margin it also certifies non-inferiority and stamps the certificate into the package.

Benchmark acceleration and quantization

After the vendor and auxiliary-asset preparation, reproduce paired eager/default/selected Runtime measurements with the included inputs and fixed checkpoint revision. Thor also requires its native backend. Keep the model and asset environments activated. For RTX 4090:

python -I -m benchmarks.regression.reproduce prepare --target rtx4090 \
  --model pi05 --mode fp8 --output pi05-inputs
python -I -m benchmarks.regression.reproduce run --prepared pi05-inputs --output pi05-results
python -I -m benchmarks.regression.serve_smoke --prepared pi05-inputs --output pi05-serving

run writes checked JSON/CSV reports and full action arrays. serve_smoke tests the actual CLI and WebSocket pipeline across two episodes. Use --mode native for default precision; FP8, numerical compilation and changed schedules are explicit selections. Reproduction guide. For additional framework comparisons, use the pinned comparison recipes.

Compare original and optimized models with instinctflash eval. Reports separate latency, action agreement and simulator task success.

instinctflash eval adapters
instinctflash eval coverage --run /path/to/run
instinctflash eval --registry plan.registry.json report --run /path/to/run

See the evaluation guide to create and run paired LIBERO / RoboTwin experiments, or benchmark details for acceleration and quantization protocols. Results: simulator screening and repeatability, checkpoints and edge latency. The expanded V2 evaluation binds latency and quality evidence to execution profiles and checks explicit control budgets. The native qualification workflow adds fresh-start admission, retained failures and checkpoint-specific evidence for each device. LingBot-VA Hub IDs retain native step counts; 2V/4A requires an explicit nfe selection. The September 9 Thor comparison separates native acceleration, FP8 Runtime gains and paired task outcomes; historical engine controls isolate additional implementation effects.

Shared BF16 fusion provides an opt-in NUMERIC path, with per-model compatibility and paired Thor regression results. Shared tensor caching and prefill separation extend native Cosmos optimization to Edge and Nano; exact caching and NUMERIC compilation remain separate options.

InstinctFlash keeps model declarations, optimization planning, runtime execution, and evidence in one inspectable path, whether it is called from Python or the command line.

A checkpoint carries a short declaration of what it is. The runtime reads the declaration, decides which optimizations are provably valid for those weights, applies them, and shows its work:

checkpoint ─▶ adapter          ─▶ planner            ─▶ engine passes        ─▶ actions
              declares what        decides what          apply and measure
              the model is         is valid (no GPU,     each optimization
                                   no weights needed)

Optimization is organized in six layers, by what each one changes:

layer changes
1 MODEL what is computed — distillation, step reduction, checkpoint compression (InstinctCompress, instinct-pdd)
2 GRAPH when work is issued — prefill extraction, CUDA-graph capture, memory planning
3 CACHE what is recomputed — KV reuse, cross-attention and episode caches
4 ATTENTION how tokens mix — FlashAttention, hybrid and linear attention
5 KERNEL how a kernel is written — backend and layout dispatch, fusion
6 HARDWARE what it executes on — fp8/int8, TensorRT, Jetson-class edge devices (serving/)

Layer 1 changes the weights and produces a checkpoint; it lives in the companion repos. Layers 2–6 change how the weights execute and produce a plan; they are the runtime in this repo. The layers are not a priority order — the runtime measures where the time actually goes and starts there.

To add your own model family, declare an instinctflash.adapters entry point and pip install your package — see examples/external_plugin/.

联系我们 contact @ memedata.com