ADR-0002: Database-Authoritative Auction Execution¶
| Field | Value |
|---|---|
| Status | Accepted |
| Date | 2026-08-29 |
| Decision owners | Dichit backend team |
| Related ADRs | ADR index |
Context¶
The unreleased auction implementation used a PostgreSQL bid-command inbox, a dedicated dispatch queue, three bid-maintenance schedulers, and a reconciliation lease. An intermediate simplification replaced lifecycle scheduling with a frequent full database sweep. The former duplicated durable state across PostgreSQL and Redis, while the latter performed unnecessary polling during normal operation. The bid inbox also made the client receive an acceptance acknowledgement before knowing whether a bid had committed.
Auctions still require correct behavior under simultaneous bids, retries, worker restarts, and overlapping workers. Client timestamps cannot safely define priority, and a realtime publication failure must not roll back a committed bid.
Decision¶
We will use PostgreSQL as the authority for bid serialization and lifecycle timing because it already owns the auction state and transaction boundary.
auction.bid.placeexecutes synchronously under an auction-rowFOR UPDATElock. Its required request id is the scoped idempotency key. The response is a terminal accepted/replayed acknowledgement or a correlated error.- The auction row stores nullable
next_transition_at, indexed with the auction id. Every lifecycle mutation recomputes or clears it transactionally. - Every committed deadline is followed by a deterministic per-auction delayed BullMQ job. A repeatable
recurring-auction-lifecycle-reconcilerqueries due rows every 30 seconds to repair missed enqueue operations. The auction queue has global concurrency one so lifecycle jobs and their ordered post-commit effects cannot overlap across worker processes. BullMQ is a wake-up mechanism, not the lifecycle source of truth. - Notifications remain queued. Realtime events are post-commit effects; clients recover from publication failure through snapshot and replay.
Consequences¶
Positive:
- Concurrent bids cannot disappear; every request commits, replays, or fails explicitly after validating the latest committed state.
- Worker restarts cannot lose lifecycle deadlines.
- Queue topology, operational alarms, and recovery behavior are substantially simpler.
Negative / costs:
- Lock acquisition order, not client arrival time, determines the serialization order of simultaneous bids.
- Normal automatic transitions depend on BullMQ delayed-job timing. A missed enqueue or lost Redis job can add up to roughly 30 seconds of recovery delay.
- Database availability is required for bidding and lifecycle progress.
Mitigations:
- Clients retry uncertain failures with the identical request id.
- Lifecycle processing is bounded, keyset-paginated, failure-isolated, and warns when lag exceeds five seconds.
- One transition per delayed job plus global queue serialization preserves observable post-end ordering through realtime publication. Immediate chains enqueue a distinct zero-delay job for each transition.
Alternatives Considered¶
Bid-command inbox and dispatch worker. Rejected because it adds delivery, cleanup, health, and replay states without improving correctness for this load; the database transaction is still required to choose the valid leader.
Individual BullMQ delayed lifecycle jobs. Rejected because Redis then holds timing authority and requires stale-version guards plus reconciliation after lost or superseded jobs.
In-process timers. Rejected because deadlines would disappear on restart and would not coordinate multiple instances safely.
Status¶
Accepted. Auction functionality is unreleased, so the implementation and its development migration are replaced without a compatibility path or backfill.