SWE Atlas 上的 Big Pickle – 代码库问答
Big Pickle on SWE Atlas – Codebase QnA

原始链接: https://github.com/PhillipChaffee/big-pickle-swe-atlas

免费隐形模型 **“big-pickle”** 在 Scale AI 的 **SWE Atlas 代码库问答基准测试**中,通过 *mini-swe-agent* 框架取得了 **50.8% 的解决率**(63/124)。 **性能亮点:** * **竞争地位:** 在 *mini-swe-agent* 类别中,big-pickle 的得分超过了所有官方排行榜条目,包括使用 Codex 框架的 GPT 模型。目前仅次于在原生 “Claude Code” 框架下运行的顶级 Claude 模型。 * **各类别表现:** 该模型在代码入门(60.7%)和架构(52.3%)方面表现最强,在 TypeScript、Python 和 Go 语言中展现出了一致的能力。 **方法论与注意事项:** * **协议:** 评估严格遵循 Scale 公布的协议,使用官方基准测试数据,并由 Claude Opus 担任裁判。 * **限制:** 此为单次试验(标准差约为 ±4.5 个百分点),并在受限的沙箱资源(4 CPU/8 GB)下运行。尽管存在这些限制,但未发生命令超时或内存溢出终止的情况。 * **模型身份:** 虽然该模型的具体来源尚未证实,但元数据表明它可能由 DeepSeek 基础设施提供支持。 本评估为自述报告,旨在提供独立的审计追踪。完整的日志和复现脚本可供社区验证。

抱歉。
相关文章

原文

Task Resolve Rate: 50.8% (63/124)big-pickle, the free stealth model on OpenCode Zen, evaluated on Scale AI's SWE Atlas Codebase QnA benchmark using the mini-swe-agent scaffold.

Run on 2026-08-11 with the official open-source harness, task data, and judge model.

Against the official SWE Atlas QnA leaderboard (updated 2026-07-28):

Model (scaffold) Task Resolve Rate
Opus 5 (Claude Code, xHigh) 63.17
Opus 4.8 (Claude Code, xHigh) 57.26
big-pickle (Mini-SWE-Agent) — this run 50.81
GLM 5.2 (Mini-SWE-Agent) 48.12
GPT-5.6-Sol (Codex, xHigh) 46.00
GPT 5.5 (Codex, xHigh) 45.43

Within the Mini-SWE-Agent scaffold class — the apples-to-apples comparison — this run outscores every entry on the official leaderboard, and it also tops the Codex-scaffold GPT entries. Only the two Claude models running on their native Claude Code scaffold score higher. Note the caveats below before treating this as a leaderboard-equivalent number.

Language Resolved Rate
TypeScript 18/31 58.1%
Python 16/29 55.2%
Go 19/38 50.0%
C 10/26 38.5%
Category Resolved Rate
Code Onboarding 17/28 60.7%
Architecture & system design 23/44 52.3%
Root-cause analysis 17/37 45.9%
Security 5/11 45.5%
API & library usage / integration 1/4 25.0%

Everything follows Scale's published protocol as closely as budget allowed:

  • Tasks: all 124 Codebase QnA tasks from scaleapi/SWE-Atlas (Apache-2.0), unmodified — including Scale's shipped mswea_qa_config.yaml agent configuration (system/instance templates, step_limit: 250).
  • Harness: Harbor v0.18.0 with Modal sandboxes, per the SWE-Atlas README.
  • Scaffold: mini-swe-agent pinned to 2.4.6 — the same minimal bash-only scaffold Scale uses for non-first-party models on the leaderboard.
  • Model: big-pickle via OpenCode Zen's OpenAI-compatible endpoint (https://opencode.ai/zen/v1), litellm route openai/big-pickle. Total consumption: 674M input / 4.3M output tokens, at $0 (the model is free during its stealth period).
  • Judge: claude-opus-4-5-20251101 — the exact judge model Scale specifies — accessed through Anthropic's OpenAI-compatible endpoint (https://api.anthropic.com/v1) with EVAL_MODEL overridden to the bare Anthropic model ID.
  • Scoring: the benchmark's own rubric-based verifier, unmodified. A task resolves only if every scored must-have rubric passes.

Read these before quoting the number:

  1. Single trial per task (-k 1). The official protocol runs 3 trials and reports the mean. At n=124, the single-trial standard error is ≈ ±4.5 points — comparable to the leaderboard's own reported error bars (±5).
  2. Reduced sandbox resources. Tasks declare 16 CPU / 16 GB; this run used 4 CPU / 8 GB to fit a personal budget. Slower command execution can only depress an agent's score (via command timeouts or OOM kills), not inflate it. Empirically it appears to have had no effect here: a scan of all 124 agent trajectories found zero command timeouts and zero exit-137 kills — no command ever hit the 900s ceiling or the memory limit.
  3. Self-reported. Scale did not run or verify this evaluation. The full per-task verifier logs in this repo allow independent auditing, and the run is reproducible from the configs here plus the public SWE-Atlas repo.
  4. Model identity unknown. big-pickle is officially unconfirmed; leaked provider errors and API response signatures suggest it is currently served by DeepSeek infrastructure. The underlying model may change without notice, so this result is a snapshot of whatever was behind the alias on 2026-08-11.
  5. Data exposure. OpenCode states that prompts to big-pickle during its free period may be used to improve the model. The benchmark's task content (already public, canary-marked by Scale) was necessarily sent to that endpoint.
  6. Two resolved tasks had unscored rubrics. On task-...ba9ad (5 of 11 rubrics) and task-...baa1d (1 rubric), the judge returned unparseable output through all 8 retries; the benchmark's verifier excludes unscored rubrics from the pass computation by design. Treating unscored-as-fail instead gives a strict-lower-bound of 61/124 = 49.2% — still above every Mini-SWE-Agent leaderboard entry. All verifier logs are included so you can apply either convention.
git clone https://github.com/scaleapi/SWE-Atlas && cd SWE-Atlas
git clone --branch v0.18.0 --depth 1 https://github.com/laude-institute/harbor.git
uv tool install ./harbor --with modal && uv tool install modal && modal setup

# from this repo: copy run_config/qa, run_config/tw, run_config/rf into
# SWE-Atlas/run_config/ (preserving the subdirectories — the scripts resolve
# .env and Scale's mswea_*_config.yaml relative to their own location),
# copy preflight.sh and .env.example into the SWE-Atlas root,
# create .env from .env.example, then:
./preflight.sh
bash run_config/qa/big-pickle_smoke.sh    # 3-task smoke test first
bash run_config/qa/big-pickle_miniswe.sh  # full 124-task run

Hard-won gotchas the configs already handle:

  • Do not pass --ak reasoning_effort with an openai/-prefixed model — Harbor silently switches mini-swe-agent to the OpenAI Responses API, which chat-completions-only endpoints like Zen don't serve.
  • Keep agent and judge credentials separate. The judge reads host OPENAI_API_KEY/OPENAI_API_BASE (via each task's [verifier.env]); the agent's Zen credentials go through --ae per-agent overrides.
  • Pass secrets to --ae as ${VAR} templates, not literals. Harbor redacts literal secrets to **** when persisting job state, which breaks harbor job resume with instant 401s. Templates round-trip and re-resolve from the host env.
  • Expect a few % of trials to die to Modal Failed to read exec stdio stream errors; harbor job resume -f <ErrorType> ... re-runs them cleanly.

Approximate cost for the full QnA run: ~$70 of Modal compute (at reduced sandbox resources; roughly 2–3× that at the declared 16 CPU/16 GB), ~$25 of Anthropic API for judging, $0 for the model.

  • results/per_task_results.csv — task ID, category, language, resolved, aggregate rubric score, rubrics passed/total
  • results/summary.json — headline numbers and breakdowns
  • results/verifier_logs/ — the judge's full per-rubric output for every task (audit trail). Notes like (flipped from raw=0) are the benchmark's own shipped verifier logic (evaluate_answer.py inverts rubrics marked negative-polarity), not post-hoc re-scoring.
  • run_config/ — the exact Harbor run scripts used (QnA smoke + full, plus untested Test Writing / Refactoring variants)
  • preflight.sh — endpoint/auth checks for both the model and the judge
  • SWE Atlas benchmark © Scale AI, Apache-2.0 — paper: arXiv:2605.08366. Per the authors' request, please treat SWE Atlas as a held-out signal of progress rather than a training target.
  • Harbor (Laude Institute) and mini-swe-agent (SWE-agent team).
  • big-pickle is served by OpenCode Zen.

Evaluation configs and results in this repo are MIT-licensed.

联系我们 contact @ memedata.com