Show HN: Jevman – AI decision models play Pac-Man

原始链接: https://opper.ai/jevman-benchmark/

GamesEach model played 100 games against the classic scripted ghosts, each until it lost all three lives. Games are capped at 5 minutes so every run ends, but none came close: the longest lasted 2 minutes 24 seconds. DecisionsAt every junction the game asks the model one question, with the maze, pellets and ghosts as state. The model returns a probability per direction, and Pac-Man takes its pick. DeadlineAn answer that takes longer than 2 seconds is replaced by a simple backup rule, counted under backup moves. ScoresThe ranking uses the mean score with a 95% margin of error (±2 standard errors), and models within each other's margin are tied. The high score is a model's best single game. ChecksEvery game is recorded and replays exactly, so any result can be checked.

相关文章

原文

Games
Each model played 100 games against the classic scripted ghosts, each until it lost all three lives. Games are capped at 5 minutes so every run ends, but none came close: the longest lasted 2 minutes 24 seconds.

Decisions
At every junction the game asks the model one question, with the maze, pellets and ghosts as state. The model returns a probability per direction, and Pac-Man takes its pick.

Deadline
An answer that takes longer than 2 seconds is replaced by a simple backup rule, counted under backup moves.

Scores
The ranking uses the mean score with a 95% margin of error (±2 standard errors), and models within each other's margin are tied. The high score is a model's best single game.

Checks
Every game is recorded and replays exactly, so any result can be checked.

联系我们 contact @ memedata.com