Skip to content
Updated Sep 12, 2026 by Barča Dvořáková · Owner: analysisactiveconceptgeneric Edit on GitHub

Notifications ​

Concept layer — frozen. The Notifications generic. Nothing here is written by a normal playbook run; a project's own feature analysis is the live document and takes every edit. This layer names no project and links to none — the dependency runs one way, from an application to its concept.

Business Context ​

Business-Level Definition ​

A notification is a deliberate act of telling a particular person that something happened which they are expected to know about or act on. It exists because the alternative is polling: a user who is not told has to keep looking, and work that nobody looked at is work that waits.

Two things distinguish it from every neighbouring capability, and confusing them is the most common way the design goes wrong at the first meeting:

  • It is addressed. A record of what happened is written for whoever later asks; a notification is written for someone, and the choice of that someone is the feature's substance. The history is Audit Log; the addressing is here. A project that has both should be able to say why an event appears in one and not the other, because the two sets are never equal — most changes are worth recording and not worth interrupting anyone about.
  • It is lossy on purpose. Its value comes from what it leaves out. A mechanism that faithfully relays everything trains its recipients to ignore it, at which point it has negative value: the organisation now believes people were told.

Everything below follows from those two sentences. The taxonomy exists so that the lossiness can be chosen rather than suffered. Preferences exist so that the recipient gets a say in it. Channels exist because "told" means different things at different urgencies, and delivery records exist because on some channels the system genuinely does not know whether the telling worked.

Requirements Definition ​

  • Some set of events, named and closed rather than "anything interesting", each with a rule saying who is told when it occurs.
  • For each recipient, a decision about whether to tell them at all, and through which channels.
  • At least one channel the system owns end to end, and any number it does not.
  • A way for a recipient to see what they have been told, and to distinguish what they have already dealt with from what they have not.
  • A way for a recipient to influence what reaches them, within limits the operator sets.
  • Delivery attempted at most as often as the event occurred, however many times the producing code runs.
  • Where a channel leaves the system, evidence of what was attempted and what happened to it.
  • Isolation: each recipient reads their own notifications and nobody else's, and where the system scopes its data, the scoping applies here as it does everywhere.

Technical Context ​

User Stories / Use Cases ​

  1. As a producing feature, I raise an event and do not decide who hears about it, so that recipient rules live in one place rather than in every feature.
  2. As a recipient, I see what I have been told, newest first, and mark items as dealt with.
  3. As a recipient, I turn down a class of notification I do not want, without turning off the ones I am relied upon to act on.
  4. As an operator, I can tell whether a message that left the system arrived, and what happened when it did not.
  5. As an operator, I can answer "was this person told?" for a specific event, weeks later.

Functional Requirements ​

The first decision: is a notification a record, or only a delivery? Every other question in this concept reads differently depending on the answer, so it is taken first and stated plainly.

  • A stored record, with deliveries hanging off it. The event and its addressing are persisted once; each channel is an attempt to surface that record. This is what makes read state, a re-readable inbox, resend, and "was this person told?" answerable at all. It costs a table that grows with events × recipients, and a retention rule to go with it.
  • Delivery only, with no record. The system renders and hands off a message, and keeps nothing but a log line. Cheap, and honest where every channel is external anyway. It costs every question above: there is no unread count, no inbox, and no way to answer what someone was told except by asking the channel provider.
  • A record derived at read time from the event history. No notification rows: the inbox is a query over events filtered by recipient rules. Nothing to keep in step and nothing to retain separately. It costs the ability to change a recipient rule without retroactively changing what people appear to have been told, and it makes per-item read state awkward, since there is no item to attach it to.

The rest of this document assumes the first, because it is the only one that supports the full requirement set — but a project whose channels are all external and whose users have no inbox is not doing it wrong by choosing the second.

The event taxonomy

  • Notifications are raised for a named, enumerated set of event types. The set is the feature's editorial policy made explicit: adding to it is a decision about someone's attention, and it should be as hard to do quietly as that implies.
  • The taxonomy is not the audit vocabulary, even where the names coincide. One is "what changed", the other is "what warrants interrupting someone". Deriving the second from the first by default — notify on every audited event, subtract exceptions — inverts the lossiness the feature exists for.

How granular the taxonomy is, is a decision, and it is not free to revisit. Preferences are keyed on it, so the taxonomy's granularity is the preference granularity — a user can never be offered a finer control than the taxonomy expresses.

  • Fine-grained — one value per distinguishable occurrence. Gives users precise control and gives templates a natural key. Costs a long preferences screen that most users will not read, and a taxonomy that grows with every feature.
  • Coarse — a handful of categories. A usable preferences screen, and few values to translate and template. Costs precision: a user who wants one of the events in a category and not another has to take both or neither.
  • Two levels — category for preferences, specific type for rendering. Usually the best of it, and the one that needs stating up front, because retrofitting a category onto a flat taxonomy means backfilling every preference and every stored row.

Whichever is chosen, adding a value later is a migration, not an addition, wherever preferences are stored as rows — see the preference decision below, which is where that cost is actually paid.

Recipient resolution

  • Every event type carries a rule naming who is told: a role on the affected record, a named assignee, participants in a conversation, the members of a group. Groups resolve to people — through whatever the project's Users & Groups capability provides — because channels deliver to people, not to sets.
  • Where the rule runs is a decision with a real consequence. Resolved by the producer, the producing feature knows its own domain and needs no notification-side knowledge of it; the cost is recipient logic scattered across features, and no single place to audit who gets told what. Resolved centrally from a generic event, the rules sit together and can be reviewed as a set; the cost is that the notification side must understand every producer's domain, which it will do imperfectly.
  • The recipient set is snapshotted, not re-derived. Whoever the rule named at the moment of the event is who was told, even if the assignment, the membership or the role changes an hour later. Re-deriving on read makes last month's inbox change contents, which destroys the one property a recipient relies on.
  • A rule that resolves to nobody is a normal outcome, not an error. A rule that resolves to hundreds is the case that decides the fan-out question below.

Boundary — the other half of "who is told". Rules attached to event types are one way a recipient set is determined; the other is a person declaring interest in a particular object. That is the Subscriptions & Watchers concept, and it is deliberately separate: subscription answers who wants to hear about this thing, this concept answers how they get told, and a system can sensibly have either without the other. Most that have both discover the same three consequences — recipients from every source are gathered, deduplicated per occurrence and only then delivered against; a subscription is never an authorisation, so entitlement is evaluated at delivery rather than at subscribe time; and a subscriber can only hear about events the system already raises, so adopting subscriptions usually means curating the event taxonomy first.

Fan-out timing — synchronous, queued, or outbox. This is the decision that determines what a missing notification means, exactly as the equivalent decision does for an audit trail.

  • Synchronously, inside the producer's transaction. Notification rows commit with the change, so a notification exists if and only if the event did. Strongest guarantee, simplest to reason about. Costs: recipient resolution and N inserts join the critical path of a business operation, so a rule resolving to a large group makes that operation slow in proportion to the group; and the notification store's availability becomes the business operation's availability.
  • Queued after commit. The producer commits and hands off. The business operation never waits for fan-out and never fails because of it. Costs the guarantee: a crash between the commit and the enqueue loses the notification silently, and "nobody was told" becomes indistinguishable from "nothing happened".
  • Post-commit, in-process. The producer commits, then calls the fan-out step directly and swallows its errors. The cheapest option and the one most projects actually ship: it has the loss semantics of the queued option — a crash between the commit and the call loses the notification — without the queue, so nothing retries and nothing records what was lost. A project choosing it says so, rather than describing it as one of the other three.
  • Transactional outbox. The intent to notify is written atomically with the change and processed asynchronously. Keeps the guarantee, keeps fan-out off the critical path, and costs another moving part plus notifications that are eventually rather than immediately visible.

The failure mode worth naming, because it recurs in every codebase that chooses the first option: a registration that is conditional on the caller supplying transaction machinery, and that returns quietly when the caller does not. It is not a missing feature — it is a silent if — and it produces exactly the outcome nobody tests for: the business action succeeds, no notification exists, no error is raised, and no log line records the decision not to notify. A project choosing synchronous fan-out states what happens when the transaction context is absent, and makes that path loud. The guard is per call site, not per feature: one producer adopting transaction hooks while the others call after commit gives the same feature two different loss semantics, and the reader of the feature page learns only one of them.

Suppression, deduplication and idempotency

  • Self-suppression — the person who caused the event is normally not told about it. Worth stating as a rule rather than leaving to each producer, and worth stating as a rule that has exceptions: confirmations ("your request was submitted") are deliberately self-addressed.
  • Within-event deduplication — one event that names the same recipient several times produces one notification. Being mentioned three times in one message is one interruption.
  • Cross-attempt idempotency — the producing path may run more than once, through a retry, a replay or a redelivered queue message. The defence is a deterministic key computed from the event rather than from the attempt, with uniqueness enforced by the store and a conflict treated as success.

What goes into that key is a decision, and both mistakes are silent. The key must identify this occurrence, for this recipient — typically the event type, the affected record, the recipient, and an identifier of the occurrence itself.

  • Omit the occurrence identifier and the key collapses to "this recipient, this record, this event type" — so the second genuine occurrence is discarded as a duplicate. A record assigned to someone, unassigned, and assigned again notifies once. Nothing fails.
  • Include the attempt, the timestamp or a generated id and the key is unique per attempt, so it deduplicates nothing while appearing to. Also nothing fails.

The occurrence identifier therefore has to come from something the producer can reproduce — the business event's own id, or the id of the row whose creation is the event. If the producer has no such handle, that is a finding about the producer, not a reason to weaken the key. The check is per producer, not per processor: a processor that accepts the handle as an opaque value cannot tell the id of a comment from a freshly generated id, so a processor that composes the key correctly proves nothing about the producers that feed it. Each event type's occurrence handle is reviewed and named in the feature documentation — a table of event type to handle — and a producer that generates one at the call site, however well argued, has chosen the second mistake. Note also that the uniqueness constraint's scope must match the key's meaning: constrained too narrowly it lets duplicates through, too widely it suppresses another tenant's unrelated event.

Preferences: what the recipient controls, and where the default lives.

Granularity — per event type × channel is the most expressive and the most work to present; per category × channel is usually what a screen can actually show; per channel only ("email me / do not email me") is defensible for a small taxonomy and nothing else.

Storage — the choice that costs the most later:

  • A row per (recipient, event type, channel), seeded when the recipient is created. Reads are trivial and the screen renders straight from storage. Costs a backfill every time the taxonomy grows and every time the recipient set does, plus a seeding path that must be idempotent and must actually be invoked on every route by which a person comes to be a recipient in a scope — which is the one that gets missed, leaving users whose preferences resolve to nothing. Where preferences are scoped, the trigger is the person's first membership of the scope, not the creation of the person.
  • Rows only where the recipient has expressed a choice, over a default table in code. Nothing to backfill, ever: a new event type simply gets its default. Costs a resolution step on every read, and a subtlety worth stating out loud — changing a default silently changes behaviour for every user who never expressed a choice, which is either the feature or the bug depending on whether it was decided.
  • Both — dense rows seeded, resolved over a default table in code. The seed is an optimisation and an audit anchor (a row exists to record who changed what); the in-code fallback is the correctness guarantee, so a seeding route that was missed leaves nobody with nothing. Costs two sources of truth for the default — the seed matrix and the in-code matrix — which are one constant, not two arrays kept in step by hand; and a behaviour change, when the in-code default changes, for exactly the users whose seed never ran.

Defaults — opt-out (on unless disabled) maximises reach and irritation; opt-in maximises quiet and guarantees that the notification nobody enabled was never seen. Most projects mix them per event type, which is the right answer and should be recorded with its reasoning rather than left as a column of booleans.

Whether every notification is optional at all is a separate question, and answering it by omission is how a project discovers that its users muted the one message that had to arrive. Some notifications are operationally or legally mandatory. A project that has such a class either excludes it from the preference surface or renders it visibly locked — silently ignoring a preference the user was allowed to set is the one unacceptable option.

Immediate, batched, or digest. Frequency is a distinct axis from channel and from event type, and a project that ships only "immediate" should know it has chosen rather than defaulted.

  • Immediate — lowest latency, highest interruption, one delivery per notification. The right default for anything action-required.
  • Batched within a short window — several notifications about the same target collapse into one. Cuts the storm where a burst of edits produces a burst of messages, at the cost of a delay on every message in the class and a bundling rule that has to decide what "the same target" means.
  • Periodic digest — everything since the last one, on a schedule. Best for low-urgency classes, and the only option that scales to a user who is a recipient of thousands of events.

Two consequences to settle before choosing anything but the first. What is the unit of read state — a digest is one delivery covering many notifications, so either the digest marks them all read (and the inbox empties for messages the user never opened) or it marks none (and the inbox repeats what the digest already said). And what happens to a notification whose target is resolved before the digest goes out: sending it anyway is noise, suppressing it makes the digest an unreliable account of the period it claims to cover.

Channels

  • A channel is a way of reaching a person. The set is enumerated and small, and each is either owned by the system or handed to a provider.
  • The distinction that matters is not "which channel" but "does the system control the endpoint". An in-application inbox is a read of stored state: the notification is delivered the moment it is stored, and success is not in question. An email, a message to a chat service, a push to a device, an SMS — each is a hand-off across a boundary, where the system knows only what the provider tells it, which is generally that the message was accepted and not that it was received.
  • A no-op provider adapter — the stub that logs and discards, useful in development and in tests — is impossible to select in production configuration. Startup refuses it; a production system quietly "delivering" to a stub is the delivery-only failure with none of the honesty.

Modelling in-app delivery as a delivery is a decision, and both answers have a cost.

  • Uniform — every channel gets a delivery record, including the in-app one. One shape, one status vocabulary, one place to ask "what happened to this notification". The cost is a class of records that no processor ever advances, which sit forever in whatever the initial status is: they are structurally indistinguishable from a genuine backlog, so every backlog alert, every queue-depth metric and every operational query has to remember to exclude them. In practice one of them eventually does not.
  • Split — deliveries exist only for channels that leave the system, and the in-app view reads the notification store directly. No dead rows, and the delivery table means exactly one thing: an outbound attempt. The cost is two shapes, and a reader who must know which channels have delivery records and which do not.
  • Uniform rows used as the visibility predicate — the in-app delivery row is written, and its existence is what the inbox list, the unread count and the bulk mark-read filter on, so a notification whose recipient disabled the in-app channel exists but is never shown. The row is not dead, it is load-bearing — while still sitting forever in its initial status and polluting every queue metric as in the uniform option. The cost is specific: the delivery table now carries two meanings, outbound queue and visibility flag, and the obvious simplification — stop writing in-app rows — empties every inbox. A project that arrives here, usually by accident, writes it down, because the next reader will propose that simplification.

Whichever is chosen, a status value that no code path ever writes is a defect worth hunting. Statuses accumulate from anticipated behaviour, then describe a lifecycle the system does not have. Every value in the vocabulary should be traceable to the code that writes it and the reader that acts on it — and any status that doubles as a reservation rather than an outcome needs its name to say so, because a value read as "this failed and will be retried" while actually meaning "a worker is holding this right now" will be misread by every operator who meets it.

Content, and the reference to what it is about

  • A notification identifies what it is about — the type and id of the affected record — and carries a way to get there. The reference tolerates the record it points at being deleted: notifications outlive their subjects, and any surface that assumes resolution will break on exactly the events that most need explaining.
  • The record type is a closed vocabulary, typed at the port. Producers are forced to use the enumerated value, not validated against it afterwards and not trusted to spell it: a free string here produces two spellings of the same entity within a month, one per producer, and every later feature that joins on the reference — subscriptions, deep-link resolution, a per-record history of who was told — inherits both and matches half the rows.
  • A follow link is not an authorisation. The recipient's access can be revoked between the notification and the click, so the target re-checks permission at the moment of the visit and fails as it would for any other unauthorised access. Equally, the notification's own text has already left the system by then — which makes what the text contains a security question, not a presentation one, and it is treated as one below.

Rendered at write time, or at read time?

  • Rendered and stored — title and body computed once, when the event happens. What the user was told is what is stored, forever, which is the only way to answer "what exactly did we say?" and the only way an external channel can be reproduced. Costs: the wording is frozen at the language and phrasing of that moment, so a corrected translation never reaches an old message, and every stored row carries whatever personal data the wording embedded.
  • Stored as a key plus parameters, rendered on read — one place to fix wording, no duplicated text, and locale resolved against the reader rather than against whoever they were when it was sent. Costs: an outbound channel has to render before hand-off anyway, so this only defers the problem for the channels that leave; and a parameter referencing a record that has since changed renders a message that is not the message that was sent.
  • Both — key and parameters retained for the in-app surface, the rendered text retained for whatever left the system. More storage, and the honest answer where evidence of external delivery matters.

Whichever is chosen, a column holds one kind of thing. A title column that holds a translation key on some rows and rendered text on others is read by a heuristic ("does it contain a dot?") that is a defect waiting for the first plain title with a full stop, and by a human who sees keys where the documentation promised text. If key and text both need storing, the column says which it holds, or they are two columns.

Read state

  • A notification is unread until the recipient says otherwise. The transition is one-way in practice — whether an unread-again affordance exists is a product decision, but read state is not a substitute for archiving or deleting, and a project that offers only "mark read" will find users using it as both.
  • Bulk marking is required at any realistic volume, and needs a cutoff, not merely "everything". Between the render and the click a new notification can arrive, and a mark-everything with no boundary silently marks it read too — the user dismisses something they never saw. The boundary is the newest item the user was actually shown.

Failure, retry, and giving up

  • For every channel that leaves the system, an attempt either succeeds, fails in a way worth retrying, or fails in a way that will never succeed. Collapsing the last two is the common error: a malformed address, a missing template, or a recipient with no address on file does not improve with time, and retrying it burns the attempt budget and the operator's attention on something only a human can fix.
  • Retry is bounded and backs off. The specific ladder is tuning, not design; what belongs in the design is that there is a bound, that exhausting it leaves a terminal, queryable state, and that reaching that state is visible to someone. A permanently failed notification that only a database query would reveal is, from the recipient's point of view, a notification that was never sent — and from the sender's, one they believe was.
  • Whether the recipient learns of the failure is a decision most projects never take. Telling them in-app that an email failed is honest and is sometimes the only way a wrong address gets fixed; it is also a message about a message, and a fresh way to be noisy.
  • Where several workers process the queue, each pending item is claimed by exactly one of them, and a claim that dies is eventually reclaimed. The mechanism belongs to the project's storage; the requirement — no double send, no permanently stuck item — belongs here.

Retention

  • Notifications and delivery records grow with events × recipients, faster than the entities they are about, and they age out of usefulness quickly: an unread notification about a record that was resolved months ago has no reader.
  • They are two separate horizons, and conflating them is a mistake. The notification is a user's inbox item and can be short-lived. The delivery record is the evidence that the system attempted to tell someone something — which is exactly what a dispute is about, and therefore a question for Legal Context below rather than for storage economics.
  • Whichever horizons apply, expiry is visible in the reading surface. An inbox that silently stops at a date invites the reader to conclude nothing happened before it.

UI/UX Design ​

Three surfaces, answering three different questions:

The signalThe inboxThe preferences
Questionis there anything new?what am I being told?what should reach me?
Shapea count, always visiblelist, newest first, unread distinguishablethe taxonomy × channels, as the user's own controls
Cost of getting it wrongignored if it is never zeroignored if it cannot be clearedignored if it cannot be understood
  • The count must be reachable. A badge that is never zero stops being read within days, so bulk clearing is a requirement rather than a convenience, and any class of notification that cannot be cleared does not belong in the count.
  • Filtering to unread is the minimum filter. Beyond that, filters are worth their complexity only once the volume justifies them.
  • The item is a link. A notification a user cannot act on from where they read it has told them about work without taking them to it.
  • Timestamps are absolute and locale-formatted on any surface that serves as a record. Relative phrasing reads well in a feed and is useless in a dispute.
  • An error state is never rendered as an empty one. "You have nothing" and "we could not load it" are different answers, and a badge that shows zero because a request failed is a lie that looks like good news.
  • Preferences state their defaults and their limits. A control that appears to be a choice but is overridden elsewhere is worse than a control that is visibly locked with a reason.
  • Accessibility: the count carries a text label rather than colour alone, the list is marked up as a list, unread is conveyed by more than colour, and each timestamp carries a machine-readable value alongside its localized text.

Internationalization & Localization ​

  • Every user-facing string — titles, bodies, event-type labels, channel names, the preference screen — resolves through the project's Translations mechanism.
  • Which locale, resolved when, is the question this feature adds. A notification is composed for someone other than the person who triggered it, so the actor's locale is never the right answer. It is the recipient's, and where the recipient has none, the scope's default, and where that is absent, a declared fallback. That chain is stated explicitly, because the failure is silent: the message goes out perfectly formed in the wrong language.
  • The choice between rendering at write time and at read time (above) decides when the locale is captured — and therefore whether a user who changes their language sees their existing inbox change with them. Both answers are defensible; only leaving it undecided is not.
  • A missing template or key for a supported locale falls back rather than failing the send, and the fallback is logged. A notification in the wrong language is a defect; a notification not sent is a worse one.

Non-Functional Requirements ​

  • Recipient isolation is absolute: a recipient reads and modifies only their own notifications and preferences, and this is enforced in the query rather than in the response.
  • Where the system scopes its data, notifications carry the scope and every read is filtered by it — see Multi-Organizations.
  • Every list read is paginated. There is no unbounded read of these tables.
  • Fan-out cost is proportional to the recipient count, so an event type whose rule can resolve to a large group is a capacity question at design time, not an incident later.
  • The notification path is not permitted to fail the business operation it accompanies, unless the project has deliberately chosen the synchronous option above and accepted exactly that.

Performance ​

  • The two hot reads are "my notifications, newest first" and "how many are unread". Both are keyed on recipient and both need index support that does not degrade as the table grows; the unread count in particular is requested on almost every page load and is the read most likely to be under-considered.
  • The queue read — "what is ready to send" — is a different access pattern again, keyed on status and due time rather than on recipient.
  • Writes are bulk inserts in proportion to the recipient set. A single event producing thousands of rows is the case that decides whether fan-out can be synchronous.

Transactional Operations ​

  • The notification and its delivery records are created together or not at all. A notification with no delivery record is one nobody will ever be sent; a delivery record with no notification has nothing to render.
  • Idempotent creation and its uniqueness constraint are one operation: check-then-insert across two statements is the race that produces the duplicate the key exists to prevent.
  • Bulk read-marking commits as a unit.
  • The relationship to the producer's transaction is the fan-out decision above, and it is the one that must be written down. Everything else here is local.
  • Producers — the features that raise events. They own the recipient rule or delegate it, per the decision above.
  • The fan-out step — resolves recipients, applies suppression and preferences, writes the record and its deliveries.
  • One dispatcher per outbound channel — claims due work, renders, hands off to a provider, records the outcome.
  • External providers — outside the trust boundary and outside the system's control. Their acceptance of a message is not evidence of its arrival.
  • A retention process — removes what has aged out, on the horizons decided above.

Diagrams & Models ​

The pipeline, with the three points where a project's answer changes the shape:

flowchart TD
    E[Business event occurs] --> R{Recipient rule}
    R -->|resolves to nobody| X[No notification. A normal outcome]
    R -->|one or more people| F{Fan-out timing}
    F -->|in the producer's transaction| W[Write with the change]
    F -->|queued or via an outbox| W
    W --> S{Suppressed?}
    S -->|actor is the recipient, or a duplicate key| X
    S -->|no| P{Preferences per channel}
    P -->|no channel enabled| X
    P -->|at least one| N[Notification record stored]
    N --> I[In-app: readable immediately]
    N --> O[Outbound channels: a delivery per channel]
    O --> D{Dispatch attempt}
    D -->|accepted by the provider| OK[Recorded as sent. Not as received]
    D -->|transient failure| RT[Retry with backoff, bounded]
    RT --> D
    D -->|permanent failure, or budget exhausted| DL[Terminal state, visible to an operator]

Two readings of this diagram are worth making explicit. The in-app branch has no failure path — that is the asymmetry the channel decision above is about. And every arrow into the "no notification" box is silent: a suppressed, unsubscribed or unaddressed event produces no error anywhere, which is why the count of notifications not sent is a metric rather than an afterthought.

API Analysis (API-A) ​

See child page: Notifications — API Analysis (API-A).

Domain Model ​

  • The notification — one row per (event occurrence, recipient), carrying the recipient, the event type, the reference to the affected record, the content or the key that renders it, the way to reach the target, read state, the deterministic idempotency key, the scope where one applies, and whatever audit columns the project gives every entity.
  • The delivery — one row per outbound attempt-track per notification, carrying the channel, the status, the attempt count, the next attempt time, the last error and any provider-side identifier. Whether in-app gets one of these is the decision above.
  • The preference — the recipient's choice for a (taxonomy value, channel) pair, stored densely or sparsely per the decision above.
  • The taxonomy of event types and the set of channels, each a closed vocabulary that producers extend deliberately.

A deletion rule worth settling explicitly: deliveries are meaningless without their notification and are removed with it, while a notification is not removed with its recipient — the record that someone was told outlives their account, unless an erasure obligation says otherwise, which Legal Context asks about below.

Logging ​

.info

  • a notification was created — event type, recipient, affected record.
  • a message was handed to a provider — channel, provider-side identifier if any.
  • read state changed — how many rows, and whether by selection or in bulk.

.warn

  • an attempt failed and will be retried — channel, attempt number, sanitized error.

.error

  • an attempt failed terminally, or failed for a reason retrying cannot fix.
  • a recipient could not be addressed on a channel they have enabled.

.debug

  • why a recipient was skipped — suppressed, preference disabled, duplicate key. This is the log line that answers "why was I not told?", which is otherwise unanswerable, because every one of those paths is a deliberate silence.

Nothing at any level carries credentials, provider secrets, or the rendered body of a message.

Monitoring ​

  • The queue's depth and its oldest pending item — the pair, not either alone, since a large queue moving quickly is healthy and a small one that is not moving is not.
  • The terminal-failure rate per channel, which is where a provider outage or an expired credential first shows.
  • The ratio of notifications created to notifications suppressed, whose drift tells you the taxonomy or the defaults have changed meaning.

If a channel's records are structurally never processed (the uniform-delivery choice above), every one of these needs to exclude them explicitly, or the first metric is permanently wrong.

Caching ​

The unread count is the one candidate, being small, hot, and requested constantly. It is also the one value where staleness is immediately visible to the user: a count that does not drop when they read something reads as broken software. Either serve it from the store, or invalidate it on every write and every read-state change — a time-based expiry is the option that looks cheapest and produces the complaint.

Indexing ​

Every access path above needs its own support: the recipient's list ordered by time, the unread subset, the idempotency key's uniqueness, the dispatcher's due-work query on channel and status and due time, the retention sweep on age, and the preference lookup by recipient. The dispatcher's index and the recipient's index have nothing in common, which is a reminder that these are two workloads in one feature.


Notifications are the one part of most systems that sends messages to people, some of whom may not be users, through infrastructure the operator does not own. What binds a given project depends on jurisdiction, sector and contract, so these are questions to settle rather than answers to copy:

Consent and the transactional line

  • Is every notification in the taxonomy transactional? Messages a person needs in order to use a service they asked for are treated very differently from messages that promote it, and the line is drawn on the message's purpose rather than on which system sent it. A taxonomy that mixes both needs to know which values fall on which side.
  • Was consent needed, and if so, is there a record of it? For channels a person must hand over an address or a device for — a mobile number, a device token, a chat account — the question is sharper than for the address they already registered with.
  • Does anything permit a recipient to be added without their action? Notifying somebody because a colleague named them, or because a group they belong to was addressed, means the system messages people who never enabled anything.

Opting out

  • Must every message carry a way to stop receiving that class? Where it must, the mechanism has to work from inside the message, for a recipient who may not be able to log in.
  • What happens when a mandatory notification meets an opt-out? Either some classes are excluded from the opt-out and that exclusion is defensible, or the opt-out wins and someone accepts that a required message can be silenced.
  • How quickly must an opt-out take effect, and does anything already queued still go out?

Retention of the record

  • Is there an obligation to prove a notice was given? If so, the delivery record is evidence and its retention is set by that obligation — not by the storage cost that would otherwise decide it. This is why the two horizons above are separate.
  • Is there a competing obligation to keep it no longer than necessary? Both pressures are normal and they point opposite ways; the resolution is a stated period, not silence.
  • Does an erasure request reach these records? Notifications name people in their content as well as in their addressing, so erasure reaches further here than the recipient column suggests.

What the content may carry

  • A message that leaves the system is an export. It crosses into a mailbox, a device or a chat service the operator does not control, cannot recall, and cannot apply its own access rules to. What may appear in the body is therefore a data-protection question, not an editorial one.
  • Does the recipient still have the right to see what the message contains? Access can be revoked between sending and reading, and the message does not revoke itself.
  • Is a bare pointer sufficient? "Something needs your attention, sign in to see it" carries no data and is often the correct answer — at the cost of a message users find useless.
  • Which providers are in the path, in which jurisdictions, and under what terms? Every external channel introduces a processor and possibly a cross-border transfer, and the provider's own logs retain message content on terms the operator does not set.
  • Are notifications sent to people who are not users — external participants, forwarded addresses? If so, the basis for holding their address, and their route to object, both need an answer that does not assume an account.

Cybersecurity Considerations ​

The content is the attack surface. Every other capability in a system decides what a person may read when they ask; this one decides what to send them unasked, to an endpoint the system does not control. The two most consequential rules follow from that:

  • Compose against the recipient's access, not the actor's. A notification is generated by one person's action and read by another, so any detail copied from the affected record must be one the recipient is entitled to see. This is the feature's characteristic information-disclosure bug: a record's name or a comment's excerpt embedded in a message sent to someone who could not open that record. It cannot be caught by the target's own permission check, because the disclosure already happened in the message. The check applies to every recipient source — producer-resolved assignees and mentions as much as subscribers — and it is applied where the recipients are gathered, once, not by whichever feature happens to add a recipient source later. A project that introduces the check for its newest source alone has documented that the older paths leak.
  • Delivery addresses are attacker-influenced input. Wherever a person can set the address a channel delivers to, that is a route to sending system-composed mail to an arbitrary destination, and wherever a person can name recipients, that is a route to using the system to message someone on their behalf. Address changes are verified, and any content the sender controls is treated as untrusted when it is rendered into a message.

Data privacy: notifications name people and reference business records by their nature, so they are personal data almost by construction. The retention and erasure questions above are the controls; the access controls are the ones below.

Authorization: reading is recipient-scoped rather than permission-scoped — the interesting rule is not which permission a caller holds but that the query can only ever return their own rows. A permission code gates the endpoint; the recipient filter gates the data, and only the second prevents one user reading another's inbox. Where the system also scopes by tenant, both apply. See Permission Model.

Rendering: message bodies embed user-supplied text. On any channel that renders markup, that text is escaped for the channel it is going to, and links in it are treated as untrusted.

Enumeration and volume: an unauthenticated or weakly authenticated trigger that causes a message to be sent is both a spam vector and a way to confirm that an address exists. Anything that can be triggered from outside is rate-limited, and its responses do not differ according to whether the address was known.


Risk Assessment ​

Business Risks ​

  • The message that did not arrive. Work waits, approvals stall, and — worse than either — the organisation believes someone was told. Every silent path in the fan-out contributes to this, which is why they are enumerated above rather than left implicit. Mitigations: a fan-out choice whose guarantee is stated, terminal failures that are visible to somebody, and a second channel for the classes that matter.
  • Fatigue. The failure mode with no error message. Users mute what they cannot control, then miss what they needed, and the mute is invisible to everyone who assumes the notification arrived. Mitigations: a taxonomy that is hard to extend casually, defaults chosen per event type rather than uniformly, suppression and deduplication, and digest options for the classes that generate volume.
  • The message that should not have been sent. Disclosure to a recipient who was not entitled to the content, or to an address that is no longer the person's. Unlike the first two, this one cannot be undone.

Technical Risks ​

  • Fan-out amplification — one event, one large group, thousands of rows and messages, inside a request. Mitigations: bound the recipient set per event, choose the fan-out timing knowing this case exists, and batch the writes.
  • Duplicate storms — a producer retried or a queue message redelivered, with a key that does not actually deduplicate. The idempotency key is the mitigation, and its composition is the thing that makes it work or silently not.
  • Provider dependence — an outage, a suspended account, an expired credential. Mitigations: bounded retry, terminal states that are queryable, alerting on the failure rate rather than on individual failures, and a design where an outbound channel's failure never degrades the in-app one.
  • Unbounded growth — two tables growing with events × recipients, faster than the entities they describe. Mitigation: retention treated as a design requirement with its horizons decided before the feature runs, not after the table is large.

Auditing, Reporting & Measurement ​

"Was this person told?" is the question this feature must be able to answer, and it is not the same as "was a notification created". A complete answer distinguishes: no notification was created (and why — no rule matched, self-suppression, a duplicate key, every channel disabled); one was created and shown in-app; one was handed to a provider; one failed terminally. A project that can report only the first and the last of those cannot explain its own behaviour to a user who asks, and that user is usually asking because something went wrong.

Whether notification events belong in the audit trail is a real question with no default answer. Sending is an action taken by the system about a person, and a preference change is a decision by a user with consequences for what they later receive — but the volume is high and most of it is uninteresting. Recording preference changes and terminal failures while leaving routine sends to the delivery records themselves is a defensible line; so is recording nothing. Recording nothing without saying so is what leaves a reader unable to distinguish a decision from an oversight.

The delivery record is the report. Status, attempt count, timestamps and provider identifier per row already answer most operational questions, and for most projects the whole reporting story is querying them. Anything aggregated is a separate build with the usual choice: against the live tables it competes with the dispatcher, against a copy it costs a copy and a lag.

Signals worth watching, each of which only this feature can produce:

  • Read rate per event type — the closest thing to a measure of whether a class of notification is worth sending. A type nobody opens is a candidate for removal, and removing one is the only action that reduces fatigue without reducing coverage.
  • Opt-out rate per event type, which is the same signal with the users' own verdict attached, and arrives earlier.
  • Suppression counts — how many recipients were skipped, and for which reason. This is the metric that makes the silent paths visible, and without it a preferences default that quietly disables an important class looks identical to an event that never occurs.
  • Time from event to read, per channel, which is the only measurement of whether the feature achieves its stated purpose. Delivery latency measures the pipeline; this measures the outcome.
  • Terminal failures per channel, and the count of recipients who have no usable address on a channel they have enabled — a population that is invisible in every other view and receives nothing.

One limit to state plainly: on every channel that leaves the system, the strongest fact available is that a provider accepted the message. Read receipts and open tracking are approximations at best, and on some channels they are unavailable or not permitted. A project should not build a report that claims to show whether people saw things when what it can show is whether messages were accepted.

Known instances ​

None recorded in this copy. Add this project's instance here as it adopts the concept — one line naming the feature that implements it. This is the only place a concept may name a project, and the template ships it empty so a fork inherits no other project's entries.