Run background workers¶
SimDB uses Celery to run long operations outside the HTTP request. When a simulation is uploaded through the REST API, the server queues the file copying and the rest of the ingestion pipeline as background tasks instead of blocking the client.
Prerequisites¶
A message broker reachable from the server and the workers. Redis is the default: see the
[celery]options inapp.cfg.
Start a worker¶
simdb_worker
Periodic tasks, including the sweep that recovers stuck ingestions, need the beat scheduler alongside the worker:
# Terminal 1: worker
simdb_worker
# Terminal 2: beat scheduler
simdb_beat
Under Docker Compose these run as the optional worker and beat services:
see Run with Docker.
Monitor tasks¶
Flower provides a web UI for queued, running, and failed tasks:
celery -A simdb.workers.celery flower --port=5555
Tasks¶
Task |
Purpose |
|---|---|
|
Copies input and output files to the server’s upload folder and updates the ingestion status. |
|
Marks a simulation as fully ingested. |
|
Runs validation checks on IMAS data. |
|
Sends email notifications. |
|
Periodic sweep that fails simulations stuck in a non-terminal ingestion state. |
Recover stuck ingestions¶
A simulation moves through non-terminal ingestion states (QUEUED, COPYING)
before reaching a terminal one (COMPLETED, COPY_FAILED,
VALIDATION_FAILED), and it can only be deleted once it reaches a terminal
state. Three mechanisms keep it from getting stuck:
A task that raises is caught, and the simulation is marked
COPY_FAILED.A task that hangs is stopped by
task_soft_time_limitand then markedCOPY_FAILED. The hardtask_time_limitis an absolute backstop.A worker that is hard-killed (SIGKILL, out of memory, node reboot) never runs its error handling, so
fail_stale_ingestions_tasksweeps up any simulation left in a non-terminal state for longer thanstale_ingestion_timeout. This requires the beat scheduler to be running.
The sweep uses the ingestion_status_updated_at column, which is refreshed on
every update, so it reflects how long a simulation has been in its current
state.
As a manual fallback, an admin can force-delete a stuck simulation whatever its ingestion state:
DELETE /v1.3/simulation/<uuid>?force=true
Run tasks synchronously¶
Setting task_always_eager = True runs tasks in-process without a broker, which
is how the test suite exercises the ingestion pipeline.