Deterministic Simulation Testing in Celld

原始链接: https://celld.dev/docs/engineering/deterministic-simulation-testing/

Hacker News new | past | comments | ask | show | jobs | submit login Deterministic Simulation Testing in Celld ( celld.dev ) 3 points by handfuloflight 2 hours ago | hide | past | favorite | discuss help Consider applying for YC's Winter 2027 batch! Applications are open till November 2. Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact Search:
相关文章

原文

We are building celld, a runtime for Cloudflare Workers and Durable Objects applications on your own machines. It can run with an S3-compatible object store as its only external service dependency.

celld is a distributed system. Making distributed systems reliable is hard, partly because we can unintentionally rely on assumptions that don’t hold in practice, even when we know the pitfalls. Peter Deutsch’s The Eight Fallacies of Distributed Computing lists eight such assumptions, including “The network is reliable” and “Latency is zero.”

A bug can depend on a particular sequence of delayed messages, failed writes, and node restarts. Those events can happen in a different order on the next test run, making the failure difficult to reproduce. We need to be able to repeat the failing run so we can investigate the cause and check whether a proposed fix really resolves the problem.

That is why we use deterministic simulation testing (DST). Our simulator is still under development and is not included in celld’s public repository, but it has already found previously unknown bugs. In this article, we’ll walk through how DST works in celld and how it helped us find, reproduce, and fix one of those bugs.

How DST works in celld

DST runs celld’s production code in an environment controlled by a simulator.

A cell in celld runs application code and has its own SQLite database. As cells do their work, celld handles events such as incoming requests, completed storage operations, and timer firings. The code that selects the next event is separate from the code that handles it. This lets the simulator control the order of events while running the same event-handling code as in production.

The simulator also controls when asynchronous tasks run and how object storage responds. It can delay a write or make it fail. It can advance simulated time without waiting for real time to pass. We use this control to test how time-dependent behavior, such as retries and alarms, interacts with other events.

We configure which requests and faults the simulator can explore. Within those limits, it uses random choices to generate requests, select what runs next, decide whether storage operations succeed or fail, and choose which faults occur and when.

A seed is the starting value for its random-number generator. With the same code, settings, and seed, we get the same sequence, intermediate states, and result. We can explore different runs by changing the seed, then repeat a failing run to trace where things went wrong.

After each action, a checker uses observed responses and stored data to test invariants: conditions we define that the system must uphold throughout a run. The test initializes celld with the simulated environment, then explores actions up to a configured limit. In pseudocode:

const simulation = createSimulation({ settings, seed });

for (let step = 0; step < settings.maxActions; step++) {
  const candidates = simulation.availableActions();
  if (candidates.length === 0) {
    checkCompletion(simulation);
    break;
  }

  const action = simulation.choose(candidates);
  simulation.run(action);
  checkInvariants(simulation);
}

Here, an action is something the simulator can advance: delivering a request, letting a task run, advancing the simulated clock, or completing a storage operation. If no actions are available, the test checks its completion conditions, such as whether all planned requests have finished.

The simplified replay below starts after Cell A has inserted a row into SQLite. Watch how the data is saved to object storage before celld receives confirmation. Other requests or background tasks can run in that gap, while celld is still waiting to return success to the client. Bugs can emerge when these operations interleave in unexpected ways. The simulator controls these steps separately so we can explore different execution orders and replay those that expose a bug.

Press Next to step through choosing an action, animating it, and revealing its result.

Candidates
  • Send the SQLite changes; let object storage receive them.
ClientcelldS3-compatible storagesimulatedCell ACell BSQLiteCell CSQLiteSQLiteWaiting for data◷ WaitingInsert rownew rowDB changesⅡ ResponseClientcelldS3-compatiblestoragesimulatedCell ACell BSQLiteCell CSQLiteSQLiteWaiting for data◷ WaitingInsert rownew rowDB changesⅡ Response
  1. Choose
  2. Run
  3. Result
Starting state

The row is in SQLite. The client is still waiting for success.

The replay above shows a successful request. Next, we’ll look at an execution order that exposed a bug in celld.

An alarm race the simulator found

An alarm schedules a cell to run application code at a specified time. The cell need not stay in memory until then: celld can unload an idle cell to make room for others and load it again in time to run its alarm.

The alarm's scheduled time is stored in the cell's SQLite database. If the cell has been unloaded, finding its alarm from SQLite alone would mean opening its database again. Doing that for every cell would be expensive, so celld instead scans wake entries in object storage. Each entry identifies a cell and when to reactivate it. Once the cell is active, celld reads the scheduled time from SQLite to determine when to run the alarm.

When we found this bug, celld reused wake entries to reduce the number of writes to object storage. For example, an application could delete a 10:00 alarm and then set a new one for 10:05. If the 10:00 wake entry was still present, celld could reuse it to reactivate the cell at 10:00, early enough for the new alarm. Once the cell was active, celld would read 10:05 from SQLite and wait until then to run the alarm.1

Deleting the 10:00 alarm clears its scheduled time from SQLite. Once that change has been saved to object storage, celld returns success to the client. A separate cleanup task removes the 10:00 wake entry later, so the client does not have to wait for that extra storage operation. The client can therefore set the 10:05 alarm while the old cleanup is still waiting to run.

After celld confirms that the 10:05 alarm is set, a wake entry must remain that can reactivate the cell by 10:05. The simulator checks that this remains true until the alarm runs or is canceled.

The figure below shows how this invariant is violated when the old cleanup runs after the 10:05 alarm has been set.

Candidates
  • Delete Cell A's 10:00 alarm.
ClientcelldS3-compatiblestoragesimulatedCell AAlarm in SQLite10:00Waiting taskDelete 10:00 wake entry10:00 wake entryPresentDelete 10:00 alarmDeletedDelete 10:00ClientcelldS3-compatiblestoragesimulatedCell AAlarm in SQLite10:00Waiting taskDelete 10:00 wake entry10:00 wake entryPresentDelete 10:00 alarmDeletedDelete 10:00
  1. Choose
  2. Run
  3. Result
  4. Check
Starting state

Cell A has a 10:00 alarm in SQLite and a 10:00 wake entry in object storage.