Transactional Outbox Integration
Trellis.EntityFrameworkCore.Outbox makes domain-event dispatch crash-safe. It captures each uncommitted domain event into an EF Core table in the same transaction as the aggregate change, then a background relay re-dispatches the events to your Trellis handlers after the commit. State and notifications commit together or not at all.
Why you need it
A plain UseDomainEvents() pipeline dispatches events after the transaction commits:
- Command handler runs, raises
OrderPlaced, repository stages the order. TransactionalCommandBehaviorcommits the transaction.DomainEventDispatchBehaviorpublishesOrderPlacedto its handlers.
If the process dies between (2) and (3) — a deploy, an OOM, a node eviction — the order is durably saved but OrderPlaced is gone. Any work the handler would have done (sending a confirmation, updating a read model, emitting an integration event) silently never happens.
The outbox closes that gap. The event is written atomically with the order, so it survives the crash and is relayed on the next poll.
The three wiring points
The outbox has a capture half (on the DbContext) and a relay half (in the service container). Both are required.
public sealed class AppDbContext : DbContext
{
// 1. Map the TrellisOutboxMessages table.
protected override void OnModelCreating(ModelBuilder modelBuilder)
{
base.OnModelCreating(modelBuilder);
modelBuilder.AddTrellisOutbox();
}
}
// 2. Add the capture interceptor where the context options are built.
builder.Services.AddDbContext<AppDbContext>(options =>
options.UseNpgsql(connectionString)
.AddTrellisInterceptors()
.AddTrellisOutboxInterceptor());
// 3. Register the relay alongside your domain-event handlers.
builder.Services.AddTrellis(trellis => trellis
.UseDomainEvents(typeof(Program).Assembly) // handlers + IDomainEventPublisher
.UseEntityFrameworkUnitOfWork<AppDbContext>()
.UseOutbox<AppDbContext>(o =>
{
o.PollInterval = TimeSpan.FromSeconds(2);
o.BatchSize = 100;
o.MaxAttempts = 10;
}));
Raise events from inside the aggregate exactly as you do today:
public sealed class Order : Aggregate<OrderId>
{
public void Place(TimeProvider clock) =>
DomainEvents.Add(new OrderPlaced(Id, clock.GetUtcNow()));
}
public sealed record OrderPlaced(OrderId OrderId, DateTimeOffset OccurredAt) : IDomainEvent;
Prefer raw DI over the builder? Call
services.AddTrellisOutbox<AppDbContext>()directly instead of.UseOutbox<AppDbContext>(). The table and interceptor wiring (steps 1–2) are identical; theUseOutboxslot simply also fails fast if you configure the outbox twice.
How it works
The capture is a SaveChangesInterceptor with a deliberate three-phase lifecycle so a failed save never loses or double-captures events:
SavingChanges— scan the change tracker for aggregates with uncommitted events, serialize each event, and add oneOutboxMessagerow per event to the currentSaveChanges. The rows enrol in the same transaction as the aggregate. The aggregate's in-memory events are not cleared yet.SavedChanges(commit succeeded) — call each aggregate'sAcceptChanges()to clear its events. Because this happens only after a successful commit, the in-pipelineDomainEventDispatchBehaviorthat runs next sees an empty event list and dispatches nothing — the relay is now the single dispatcher.SaveChangesFailed— detach the outbox rows the interceptor staged. The aggregate keeps its events, so a retry re-captures cleanly without leaving orphaned rows.
The relay (OutboxRelay<TContext>) is a hosted BackgroundService. Each poll it opens a bookkeeping scope, reads a batch of pending rows ordered by the monotonic Sequence column, rehydrates each event from its stored type name, and publishes it through IDomainEventPublisher (the same fan-out the pipeline would use) in its own per-message scope — so a handler that injects TContext gets a fresh context, not the relay's bookkeeping one, and its tracked changes never ride the relay's save. It then marks each row processed and saves the batch.
Delivery semantics
The guarantee is at-least-once delivery, and delivery means every handler completed:
- A domain message is marked processed only once every registered handler has completed. A handler that throws leaves the message pending, and the retry re-invokes only the failed handlers —
OutboxMessage.CompletedHandlersrecords the ones that already succeeded, so a successful handler's side effect is never repeated because an unrelated sibling failed. (In-pipeline dispatch still swallows, because it runs post-commit and has nothing to retry with.) - Integration events produced by handlers that did succeed are staged even on a failed attempt — they must be, since the retry skips those handlers.
- Resolution, deserialization, and publisher failures count toward
MaxAttemptswhen failure bookkeeping persists. A bookkeeping save failure instead fails the drain: it logsDrainFailed, waitsPollInterval, and leaves the previous lease/progress in the database. Repeated bookkeeping failures are not bounded by the message attempt cap; monitor them separately. - A crash between dispatch and the relay's bookkeeping save re-delivers the message to every handler. Make your handlers idempotent. The
OutboxMessage.Id(a UUIDv7) is a stable per-message key you can use for consumer-side de-duplication.
Upgrading? This replaces the previous behavior, under which a throwing handler was silently treated as delivered. Messages whose handlers fail now retry and can park — alert on
OutboxRelay.MessageParked— and theTrellisOutboxMessagestable gains aCompletedHandlerscolumn, so add a migration before deploying.
Retries, dead-lettering, and replay
When a message fails for an infrastructure reason, the relay does not retry it immediately — that would spin a poison message (or one caught in a brief outage) in a tight claim-fail-retry loop that hammers the database. Instead, each failed attempt leases the row forward with an exponential backoff: the wait after the nth attempt is RetryBackoff × 2ⁿ⁻¹, capped at MaxRetryBackoff. A transient blip is ridden out by a slightly later retry, while a persistently failing message backs off to a slow, steady cadence.
The backoff also carries a small per-message jitter that only ever shortens the wait. It is derived deterministically from the message id — no randomness, so behavior is reproducible in tests — and it spreads the retry times of messages that failed together, so when a downed dependency recovers they don't all retry in the same instant and flood it.
After MaxAttempts failures the message is dead-lettered (the log calls it parked): it stays in TrellisOutboxMessages with ProcessedAt still null, the scan skips it so it never blocks later messages, and the relay logs OutboxRelay.MessageParked at Error — the signal to alert on. The row is kept for inspection; it is a soft dead-letter, not a separate queue.
Tuning the retry behavior
Every knob has a sensible default, so set only the ones you want to change:
.UseOutbox<AppDbContext>(o =>
{
o.MaxAttempts = 10; // attempts before dead-lettering
o.RetryBackoff = TimeSpan.FromSeconds(30); // first backoff; doubles each attempt
o.MaxRetryBackoff = TimeSpan.FromHours(1); // ceiling on the per-retry wait
o.RetryBackoffJitter = 0.5; // 0 disables; 0.5 = up to 50% earlier
});
With the defaults a message keeps retrying for roughly a few hours before it dead-letters.
Replaying dead-lettered messages
Once you have fixed the cause — deployed the missing handler assembly, corrected a bad payload — re-drive the dead-lettered messages by resolving IOutboxMaintenance from a scope (an admin endpoint, a CLI, or a maintenance job):
app.MapGet("/outbox/dead-lettered", (IOutboxMaintenance outbox, CancellationToken ct) =>
outbox.GetDeadLetteredAsync(cancellationToken: ct));
app.MapPost("/outbox/dead-lettered/{id:guid}/replay", (Guid id, IOutboxMaintenance outbox, CancellationToken ct) =>
outbox.ReplayAsync(id, ct));
app.MapPost("/outbox/dead-lettered/replay-all", (IOutboxMaintenance outbox, CancellationToken ct) =>
outbox.ReplayAllAsync(ct));
Replaying resets a message's attempt count and clears its lease so the relay drains it again on the next poll; ReplayAsync / ReplayAllAsync return how many rows they re-drove. It is safe to run while the relay is live — the relay never touches a dead-lettered row, so there is no race — and it re-runs the handlers, so, as always, keep them idempotent.
Replay covers only dead-lettered messages (those that exhausted
MaxAttempts). It does not re-emit successfully-processed messages — the outbox is a delivery buffer, not a replayable event log. To re-emit already-delivered events (to rebuild a read model or onboard a new subscriber), replay from your source of truth or from a broker's retained log.
Domain events vs. integration events
A domain event is internal to your bounded context — raised by an aggregate, dispatched in-process, free to speak the domain's ubiquitous language. An integration event is the stable, versioned contract you publish to the outside world. Relaying raw domain events externally couples other systems to your internal model; the outbox lets you keep them separate and publish a deliberate contract instead.
Model the contract as an IIntegrationEvent (primitive/nullable members, no internal value objects), then translate from the domain event with an ordinary domain-event handler that adds to the scoped IIntegrationEventCollector:
// The external contract.
public sealed record OrderPlacedIntegrationEvent(Guid OrderId, string CustomerEmail, decimal Total, DateTimeOffset OccurredAt)
: IIntegrationEvent;
// The translator: a domain-event handler that emits the contract.
public sealed class OrderPlacedTranslator(IIntegrationEventCollector collector) : IDomainEventHandler<OrderPlaced>
{
public ValueTask HandleAsync(OrderPlaced domainEvent, CancellationToken cancellationToken)
{
collector.Add(new OrderPlacedIntegrationEvent(
domainEvent.OrderId.Value, domainEvent.CustomerEmail.Value, domainEvent.Total.Amount, domainEvent.OccurredAt));
return ValueTask.CompletedTask;
}
}
// Wire the consumer side (or swap the publisher for a broker adapter).
builder.Services.AddTrellis(trellis => trellis
.UseDomainEvents(typeof(Program).Assembly) // translators are domain-event handlers
.UseIntegrationEvents(typeof(Program).Assembly) // publisher + collector + in-process consumers
.UseEntityFrameworkUnitOfWork<AppDbContext>()
.UseOutbox<AppDbContext>());
How delivery flows. The relay opens IIntegrationEventCollector.BeginTranslation() around domain publishing and draining. All drained events become new integration rows in the same save as partial handler progress, even if a sibling failed or a translator threw after adding an event. A later drain publishes them through IIntegrationEventPublisher. The source domain change is durable, but the domain row need not yet be fully processed. Completed translators are skipped on retries; failed translators can create fresh rows with new IDs, requiring business-identity deduplication. Add outside an active relay lease throws instead of silently losing events.
Startup validation. Before its background loop starts, the relay resolves IReportingDomainEventPublisher in a fresh scope. Integration-enabled hosts must also resolve IIntegrationEventPublisher; domain-only hosts need none. This validates DI construction, not broker connectivity. Post-commit Trellis ETag synchronization and event clearing ignore a newly canceled token because persistence already succeeded.
The default IIntegrationEventPublisher fans out to in-process IIntegrationEventHandler<T> registrations — ideal for a modular monolith and for tests. To deliver to other services, replace that one registration with a message-broker adapter; aggregates, translators, and the outbox are unchanged. Delivery is at-least-once and a retried domain event re-runs its translator, so a consumer may see the same integration event more than once (with a different OutboxMessage.Id each time) — dedupe on business identity, not on the message id.
Writing a broker adapter
Two things separate an adapter that works from one that quietly loses the inbox's guarantee.
Publish through the message, not the event. IIntegrationEventPublisher has exactly one method, PublishAsync(OutboundIntegrationMessage, CancellationToken), whose MessageId is the outbox row's own id. Stamp it on the wire (ServiceBusMessage.MessageId, a Kafka header, and so on) so the consumer's (ConsumerId, MessageId) inbox dedup can collapse redeliveries of that row. Minting a fresh id per publish attempt makes every redelivery look like a new message and defeats the inbox. The bare-event overload was deliberately removed so an adapter cannot publish without the id by accident.
Identify the event by a logical name, not by its CLR type. The outbox stores Type.AssemblyQualifiedName, which is correct for its own in-process relaying and unusable across services: the consumer's assemblies differ, and the string embeds an assembly version. Annotate contracts and resolve them through IntegrationEventNameMap:
[IntegrationEventName("orders.order-placed.v1")]
public sealed record OrderPlacedIntegrationEvent(Guid OrderId, DateTimeOffset OccurredAt) : IIntegrationEvent;
// Producer and consumer each build the map over their own contract assembly.
var names = IntegrationEventNameMap.FromAssemblies(typeof(OrderPlacedIntegrationEvent).Assembly);
Include a version segment in the name. Once a message carrying it sits on a queue or in another team's code, changing the name is a breaking change; a new version can be consumed side-by-side instead.
On the receiving end, build an IntegrationEnvelope from the wire id and the resolved type and hand it to IInboxDispatcher. Acknowledge the broker message on both Processed and SkippedDuplicate — the dispatcher contract says both mean the message is durably accounted for.
Bear in mind what the message id can and cannot do: it collapses redeliveries of a single outbox row, but a retried domain row re-runs its translator and stages a genuinely new row with its own id. That second duplicate still needs business-identity deduplication.
Persist-on-failure (FailAfterCommit) events
TransactionalCommandBehavior commits a Result.FailAfterCommit outcome (persist-on-failure), but DomainEventDispatchBehavior does not dispatch domain events for any failed result — by design those events are discarded, not a durable buffer.
The capture interceptor cannot see the command result, so with the outbox enabled it captures events from every commit, including persist-on-failure commits, and the relay delivers them. Enabling the outbox therefore changes this behavior: events raised on a persist-on-failure path become durable and dispatched. If you rely on the base suppression, do not raise domain events on persist-on-failure paths — return a success result for the events you want delivered, and model post-failure side effects explicitly (a follow-up command, or a dedicated outbox row).
Serialization constraints
Events are serialized and rehydrated with a Trellis-owned System.Text.Json options instance that adds a Maybe<T> converter:
- Round-trips: value objects that carry a
[JsonConverter]attribute (the Trellis scalar and composite primitives) — the converter travels with the type.Maybe<T>members also round-trip: a present value serializes as the underlying value and an absent one as JSONnull. - Does not round-trip: converters registered only through a caller-supplied
JsonSerializerOptionsfactory — they are not consulted by the outbox serializer. Shape those members with attribute-driven value objects or a nullable transport (string?,decimal?, …).
Operating the outbox
- Alert on dead-lettered messages. Page on the
OutboxRelay.MessageParkederror log (or on rows whereProcessedAt IS NULL AND Attempts >= MaxAttempts); those exhausted their retries and need a fix plus a replay. Transient per-attempt failures log at Warning (OutboxRelay.RelayAttemptFailed) and self-heal via the backoff, so they should not page on their own. - Prune processed rows. Rows with a non-null
ProcessedAtare a spent delivery buffer — a periodic job can delete old ones with no loss of source-of-truth state. The aggregate tables remain authoritative. - Keep the producing assemblies loaded. The relay resolves each event by its assembly-qualified type name, so the worker process must reference the assemblies that declare your events.
- Run as many relay instances as you need. Each drain atomically claims a batch with a lease (
LockedBy+LockedUntil), andLockedByis an optimistic concurrency token, so concurrent instances never both own a row's lease at once, and a drain that outlived its lease abandons its bookkeeping write — loggingOutboxRelay.LeaseLost— rather than clobber the instance that reclaimed the row. No leader election or distributed lock is needed. Delivery is still at-least-once — a crash between publish and the relay'sSaveChanges, or a batch that outlives its lease, can re-deliver — so keep handlers idempotent. SetLeaseDurationcomfortably above the worst-case batch publish time and keep node clocks reasonably in sync, since the lease is compared against each relay's wall clock. - One outbox per composition.
UseOutbox<TContext>()throws if called twice. Multiple relays in one process are not supported by the builder slot today.
Outbox vs. event sourcing
The outbox is not event sourcing. The TrellisOutboxMessages rows are a transient messaging buffer: they exist only until they are relayed and may be pruned afterward. Your aggregate tables remain the single source of truth, and you never rebuild state by replaying outbox rows. Event sourcing, by contrast, makes the event log itself the source of truth. Reach for the outbox when you want reliable delivery of notifications about state changes; reach for event sourcing when the events are the state.
Related guides
- Entity Framework Core Integration — repositories, unit of work, and the interceptors the outbox builds on.
- Mediator Pipeline — domain-event dispatch, the
IDomainEventPublisherseam, and pipeline ordering.