The transactional outbox pattern stores a business change and its event in the same database transaction. A separate relay publishes the event to Kafka, allowing delivery to resume after a failure without losing the intent to notify other services.

Consider a Spring Boot interview-booking service backed by PostgreSQL. Confirming a booking should also notify another service through Kafka. This guide follows that operation from database commit to consumer processing, with illustrative SQL and Java fragments. It connects the backend and messaging technologies in my portfolio to a practical reliability problem: keeping business state and downstream events consistent.

Why a database write and a Kafka send can disagree

A straightforward implementation inserts a booking and then sends a BookingConfirmed event. If PostgreSQL commits but the application stops before publishing, the booking exists without a notification. Reversing the operations creates the opposite possibility: a consumer sees an event for a booking that never commits. Retrying the whole HTTP request introduces another question: has the original booking already succeeded?

The AWS transactional outbox guidance describes this dual-write problem and the need to handle duplicate messages. The pattern changes the durable unit of work. Instead of requiring the database and broker to become available at exactly the same moment, the application records the booking and an event awaiting publication together.

Kafka transactions are useful within their supported boundaries, but they do not automatically make a PostgreSQL commit atomic with a Kafka commit. Spring Kafka's transaction documentation explicitly discusses failure of a synchronized transaction after the primary transaction has committed. An annotation around two calls is not evidence that the dual-write gap has disappeared.

Design a PostgreSQL outbox table

The following schema is an illustrative starting point for a polling relay. Each row describes an immutable event. Publication status belongs to delivery bookkeeping, while the event identifier and payload remain stable across retries.

CREATE TABLE outbox_event (
    event_id uuid PRIMARY KEY,
    aggregate_id uuid NOT NULL,
    event_type text NOT NULL,
    payload jsonb NOT NULL,
    created_at timestamptz NOT NULL DEFAULT now(),
    published_at timestamptz
);

CREATE INDEX outbox_event_pending_idx
    ON outbox_event (created_at, event_id)
    WHERE published_at IS NULL;

Here, aggregate_id identifies the booking and event_id identifies this particular event. A later cancellation gets another event identifier while retaining the booking identifier. Add an explicit payload version so consumers can distinguish intentional contract changes from malformed messages. Generate the event identifier once when recording the event, rather than creating a new identifier for every publishing attempt.

Keep the payload focused on what consumers require. A booking identifier, event version and relevant state may be enough; copying a complete user profile into every event unnecessarily expands access and retention. Define whether the event contains an immutable fact or asks the consumer to fetch current state. Those meanings differ when a booking changes before the consumer catches up.

Write the booking and event in one Spring transaction

Both repository operations must participate in the same PostgreSQL transaction. The following service-method fragment assumes that a bean named transactionManager manages that database and that both repositories use it. Entity mappings, validation and repository implementations are omitted.

@Transactional(
    transactionManager = "transactionManager",
    rollbackFor = Exception.class
)
public UUID confirmBooking(ConfirmBooking command) {
    Booking booking = bookings.reserve(command);
    UUID eventId = UUID.randomUUID();

    outbox.insert(
        eventId,
        booking.id(),
        "BookingConfirmed",
        bookingEventPayload(booking, 1)
    );
    return booking.id();
}

Invoke the method through the Spring-managed service proxy. The Spring transaction reference explains that self-invocation does not trigger transactional interception in the default proxy mode. Check propagation settings as well: an outbox repository using an independent transaction can undermine the intended all-or-nothing operation.

The reservation still needs its own concurrency invariant, enforced by appropriate database constraints or locking. Outbox does not prevent double booking. Similarly, an HTTP idempotency key should identify a retried business request and be persisted with its outcome. Event deduplication and request deduplication solve different problems, even when both use unique identifiers.

Publish with a relay that expects interruptions

A polling relay selects pending rows, publishes them and records successful publication. PostgreSQL supports FOR UPDATE SKIP LOCKED for queue-like workloads with competing workers, as described in its SELECT documentation. The lock exists only for the database transaction's lifetime.

BEGIN;

SELECT event_id, aggregate_id, event_type, payload
FROM outbox_event
WHERE published_at IS NULL
ORDER BY created_at, event_id
LIMIT 20
FOR UPDATE SKIP LOCKED;

-- Relay code publishes selected events and awaits broker acknowledgements.
-- It updates published_at only for acknowledged events, then commits.
-- A failure rolls back uncommitted bookkeeping; retries reuse event_id.

This is a transaction-flow sketch, not a complete worker. Holding locks while waiting for Kafka consumes a database connection, so use small batches, bounded publish waits and transaction timeouts. Returning from an asynchronous send call is not a broker acknowledgement. For higher throughput, consider a short claim transaction with durable leases, ownership checks and recovery for expired claims; releasing locks before publishing without such a mechanism invites competing workers.

A crash after Kafka accepts a message but before published_at commits leaves the event pending. Publishing it again is the safe recovery behavior. Configure durability appropriate to the workload, retry transient failures with backoff, and make persistent failures visible. Never mark an event published merely because the retry budget was exhausted.

When to use Debezium instead of polling

A change-data-capture relay can read committed outbox inserts from PostgreSQL's replication stream. The Debezium Outbox Event Router maps an outbox row into a Kafka message, including an event identifier and aggregate key. Its default column names differ from the polling schema above; configure explicit mappings or adopt the documented schema.

Choose one publication mechanism for a given event flow. Running an independent polling publisher alongside CDC without coordination creates another duplicate source. With CDC, table cleanup also needs a retention policy that respects connector progress and recovery requirements; a polling-style published_at field is not automatically maintained by the connector.

CDC shifts operational responsibility rather than removing it. The PostgreSQL connector documentation covers replication slots, snapshots and WAL retention. Track retained WAL and connector lag so a stopped relay cannot quietly consume database storage. A small service may prefer polling initially; an organization already operating CDC may find a connector easier to maintain.

Make the consumer's database effect idempotent

For a consumer that updates a local read model, store a processed-event marker in the same transaction as the business update. A unique key on (consumer_name, event_id) lets different logical consumers process the same event independently while suppressing duplicates within one consumer.

CREATE TABLE processed_event (
    consumer_name text NOT NULL,
    event_id uuid NOT NULL,
    processed_at timestamptz NOT NULL DEFAULT now(),
    PRIMARY KEY (consumer_name, event_id)
);

-- Execute inside the transaction that updates the local read model.
INSERT INTO processed_event (consumer_name, event_id)
VALUES ('booking-read-model', :event_id)
ON CONFLICT DO NOTHING
RETURNING event_id;

:event_id represents a bound application parameter. If the insert returns a row, perform the business update; otherwise skip that already-processed event. Commit both before acknowledging the Kafka offset. If processing fails, roll back the marker too. PostgreSQL documents ON CONFLICT and RETURNING in its INSERT reference.

A crash after the database commit but before the offset commit causes redelivery; the marker makes that safe for this database effect. Retain markers for the supported replay window. An external email or payment call cannot join this local transaction: use the provider's idempotency contract, another durable work queue, or reconciliation suited to that side effect.

Define ordering at the business boundary

Use the booking identifier as the Kafka message key when events for one booking belong together. Kafka preserves partition order, but it cannot repair events that a concurrent relay published in the wrong business order. Its delivery-semantics documentation also distinguishes Kafka transaction guarantees from effects in external systems.

The polling sketch sorts by creation time for repeatable selection, not a guarantee of business or commit order. Another worker can skip a locked confirmation and publish its cancellation first. If order matters, serialize publication per aggregate, or assign an aggregate version transactionally and make consumers detect gaps and stale updates. Choose the mechanism before allowing concurrent publication; deduplication alone does not solve ordering.

Test failures at the commit boundaries

Failure cases for a booking outbox implementation
Injected failureExpected recoveryEvidence
Outbox insert failsBooking transaction rolls backNo partial booking remains
Kafka is unavailableCommitted events remain pendingBacklog drains after recovery
Relay stops after sendEvent may be delivered againStable event ID; one consumer database effect
Consumer stops after commitRedelivery is deduplicatedMarker and business update agree
Two events raceChosen ordering policy appliesNo cancellation overwritten by stale confirmation

Track the age of the oldest pending event alongside backlog size, retry counts and consumer processing delay. Zero publication errors can still hide a worker that has stopped entirely. Connect event IDs to traces without logging entire payloads, following the signal-correlation approach in the OpenTelemetry guide. Cleanup should preserve the agreed replay and investigation window.

Common transactional outbox questions

Does transactional outbox provide exactly-once delivery?

It makes the business write and event record atomic within one database. The relay can deliver duplicates, so consumers need idempotent handling. Describe the specific business effect you protect rather than promising universal exactly-once execution.

Can Redis locking replace the outbox?

A lock coordinates competing operations; it does not preserve an event after the application crashes between a database commit and a broker send. Redis may still serve other purposes, such as the read optimization discussed in the Spring Boot caching guide.

When is an outbox unnecessary?

If all required work fits within one database transaction and there is no downstream publication requirement, the extra relay may add little value. Use an outbox when losing an event after a successful business commit is unacceptable and the team can operate eventual delivery, retries and recovery.