Kimi K3 现已接入 Telnyx Inference API
Kimi K3 Now Available via Telnyx Inference API

原始链接: https://telnyx.com/release-notes/kimi-k3-telnyx-inference

Moonshot AI 的 **Kimi K3** 是全球首个三万亿参数级别(2.8T 参数)的开源模型,现已通过 Telnyx 推理 API 提供服务。 Kimi K3 专为高性能代理工作流设计,具备 100 万 token 的上下文窗口、原生多模态能力(支持文本、图像和视频)以及可配置的推理等级。它支持工具调用、JSON 结构化输出和自动前缀缓存等高级功能。 Kimi K3 弥合了开源模型与闭源前沿模型之间的差距,其性能足以与 OpenAI 和 Anthropic 的顶尖模型相媲美。通过利用 Telnyx 自有的 GPU 基础设施和兼容 OpenAI 的 API,开发者可以轻松地将该模型集成到现有技术栈中。 **核心规格:** * **架构:** Kimi Delta Attention 和 Attention Residuals。 * **定价:** 每百万 token 0.27 美元(缓存输入)、2.70 美元(输入)和 13.50 美元(输出)。 此次发布标志着人工智能领域的重大转变,证明了开源权重模型在强大的可扩展基础设施支持下,已能够媲美专有的前沿模型。

月之暗面(Moonshot AI)已发布其新款 Kimi K3 模型的开放权重,现可通过 Telnyx Inference API 使用。 Telnyx 强调了使用其平台的几项优势:通过利用其在美国、欧盟、亚太和中东及北非地区自有并运营的 GPU 基础设施,他们最大限度地降低了延迟并消除了跨服务商跳转。他们强调了隐私保护,指出其采取零数据留存政策,不会存储任何提示词或生成内容。此外,通过掌控硬件,Telnyx 维持了极具竞争力的定价:输入 Token 每百万 2.70 美元,输出 Token 每百万 13.50 美元,并默认开启提示词缓存(Prompt Caching)。 K3 模型被定位为前沿级模型,在编程和智能体任务方面表现尤为出色。开发者可以通过兼容 OpenAI 的端点访问 Kimi K3,从而简化集成过程。 尽管该公告引发了人们对模型性能的关注,但 Hacker News 上的讨论也涉及了用户对 Telnyx 强制性 KYC(了解你的客户)要求的担忧,一些用户认为这对于访问服务而言过于侵入隐私。
相关文章

原文

Kimi K3, Moonshot AI's 2.8-trillion-parameter flagship model, is now available on the Telnyx Inference API. It is the world's first open-source model in the 3-trillion-parameter class, built on Kimi Delta Attention and Attention Residuals with a 1M-token context window and native vision capabilities.

What's new

  • New model available: Kimi K3 (model ID: moonshotai/Kimi-K3) is now selectable on the Telnyx Inference API alongside existing models including Kimi K2.6, GLM-5.2-FP8, and MiniMax M3.
  • 2.8T parameters: The largest open-weight model available on Telnyx Inference. First open-source model to reach the 3-trillion-parameter class.
  • 1M token context window: Supports codebase analysis, long document processing, and multi-turn agent sessions with stable long-context performance.
  • Native vision: Accepts text, images, and video input within the same model. Multimodal reasoning without a separate vision adapter.
  • Configurable reasoning effort: Three levels (low, high, max) to trade compute for depth of reasoning per request.
  • Tool calling and structured output: Supports function calling, dynamic tool loading, and JSON schema constrained output for agentic workflows.
  • Prompt caching by default: Automatic prefix caching for repeated prompt prefixes across requests.

Why it matters

The competitive advantage in AI is shifting from who builds the smartest model to who builds the infrastructure that decides where every request runs, and K3 is evidence that the model side of that equation is solving itself. Kimi K3 is the first open-source model to reach 2.8 trillion parameters, and on benchmarks for coding, reasoning, and agentic knowledge work, it competes with closed-source frontier models from Anthropic and OpenAI, proving that open-source is not far behind the frontier labs, and in some cases is already there.

K3 now runs on Telnyx-owned GPU infrastructure and can be access via the OpenAI-compatible API.

Pricing

Token TypePrice per 1M tokens
Cached Input$0.27
Input$2.70
Output$13.50

Learn more in the Inference documentation or try it in Mission Control.

联系我们 contact @ memedata.com