Skip to content
Updated Sep 12, 2026 by Barča Dvořáková · Owner: analysisactiveconceptgeneric Edit on GitHub

Comments ​

Concept layer — frozen. The Comments generic. Nothing here is written by a normal playbook run; a project's own feature analysis is the live document and takes every edit. This layer names no project and links to none — the dependency runs one way, from an application to its concept.

Business Context ​

Business-Level Definition ​

Comments let people write to each other about a particular record, in the place where that record is worked on, rather than in a channel somewhere else. The value is not the writing — every organisation already has somewhere to write. The value is the attachment: the discussion arrives with the thing it is about, for whoever opens it next, including people who were not there at the time.

That is also the reason it is one shared capability rather than a field on each record that needs one. A per-record "notes" column is easy and gives up almost everything: no author, no ordering, no addressing anyone, no history, and a different behaviour on every screen. Once discussion exists in more than one place, the question "what was said about this?" needs a single answer.

Two consequences follow immediately and are worth stating in business terms before any design. Comments become, in practice, the record of why — the reasoning that no structured field captured. And they are the least structured store the system has: freeform text, written quickly, about anything, by anyone who can see the record. Everything difficult about this capability — retention, erasure, disclosure, moderation, search — follows from those two facts together.

Requirements Definition ​

  • Discussion attaches to a record and is read in that record's context.
  • One mechanism serves every kind of record that needs it, including kinds that do not exist yet.
  • Each comment says who wrote it and when, and that attribution survives the author being renamed, deactivated or removed.
  • People can be addressed directly, so a comment can ask someone specific for something.
  • Authority over a comment derives from authority over what it is attached to; a comment is never more visible than the record it discusses.
  • What happens when a comment is edited or removed is defined and observable, rather than left to whichever code path ran.
  • The stream is readable at any size, which means paginated, ordered deterministically, and bounded in what a single read can cost.

Technical Context ​

User Stories / Use Cases ​

  1. As someone who can read a record, I read the discussion attached to it, oldest or newest first, without holding any comment-specific grant.
  2. As a participant, I write a comment on a record I can read, and address specific people or teams in it.
  3. As an author, I correct what I wrote, and what readers see afterwards is the correction.
  4. As an author, I remove what I wrote, and readers can tell that something was removed rather than that it never existed.
  5. As a moderator, I remove someone else's comment, and that act is distinguishable from the author removing their own.
  6. As someone addressed in a comment, I learn about it without opening the record.
  7. As an investigator or a support agent, I read what a comment used to say, subject to a grant that ordinary readers do not hold.
  8. As a reader of a busy record, I find the comment I remember without scrolling through hundreds.

Functional Requirements ​

What the comment is attached to

The attachment model is the first decision, and everything downstream inherits it.

  • One polymorphic reference — a target kind plus a target id. One table, one read surface, one component, and a record kind that does not exist yet is already supported. The costs are real and permanent: no foreign key, so references can dangle and every reader tolerates a comment whose target no longer resolves; no cascade, so removing a target leaves its discussion behind unless something removes it deliberately; and authorisation cannot be expressed in the schema, because "who may read this" now depends on a value in a column.
  • A table per commentable record kind. Foreign keys, cascades, and authorisation that reads naturally. The cost is a new table, a new read surface and a new component per kind, and the guarantee that they will drift — the fourth one will support something the first three do not.
  • An intermediate discussion entity that records point at. The record owns a reference to its discussion, and comments belong to the discussion rather than to the record. Restores the foreign key and gives a natural home for anything that belongs to the conversation rather than to a message — a status, a participant set, a subscription list. Costs one more join on every read and a row that must be created before the first comment, which is a real ordering problem under concurrency.

A project that picks the polymorphic reference owes two things it will otherwise discover late: a declared, closed set of supported target kinds — not "any string" — and a stated answer for what happens to comments when the target is removed. The honest answers are few: removed with the target in the same write; retained and still readable through some other surface; or retained but unreachable — the rows stay, and every read path answers as it would for a target that does not exist, because the target check runs before the comment query. The last is a complete answer and a common one; it only becomes a defect when nobody wrote it down and a reader files the missing discussion as data loss.

A second, independent fork: does a comment attach to the record, or to a place inside it?

Attaching to the record is what most systems mean, and it is stable: the target has an identity that does not move.

Attaching to a region — a text selection, a page area, a row, a field — is a materially different capability wearing the same name. The anchor must survive the content changing underneath it, which needs either an anchoring scheme that can be repaired or an accepted state where a comment is orphaned: still real, no longer pointing anywhere. A project that wants anchored comments should treat them as their own analysis and not assume the record-level design stretches to cover them; a project that does not want them should say so, because it is the first thing someone will ask for.

Conversation shape

How deep may a conversation go?

  • Flat. Every comment is a sibling, ordered by time. Trivially pageable, trivially rendered, and honest about what most discussions actually are. It cannot express "this is a reply to that", so on a busy record two conversations interleave and neither is followable.
  • Bounded depth — replies, but only to a top-level comment. Buys the reply relation and keeps every property that makes flat easy: the read is at most two queries, the render is a list of lists, and pagination has an obvious unit (the top-level comment and its replies). The bound is arbitrary and will occasionally annoy someone, which is the price of it being enforceable.
  • Unbounded nesting. Expressive, and it moves the cost everywhere else: the read becomes recursive or needs a materialised path, deep threads become unreadable at the third indent, and pagination stops having a well-defined unit — a page of what?

Whichever is chosen, two invariants hold. A reply belongs to the same target as the comment it replies to, checked on write rather than assumed. And depth is enforced on write, not on render: a rule applied only by the client is not a rule.

Is a thread a first-class object, or just a parent pointer?

This looks like a modelling detail and is actually a product decision. A first-class thread can carry state — most usefully open / resolved — and that single piece of state changes what the capability is for: from a discussion log into a lightweight worklist attached to a record. It also gives subscriptions, participant lists and per-thread muting somewhere to live.

The cost is not the row. It is that "resolved" is a claim someone can disagree with, so a project now needs answers for who may resolve, whether a resolved thread reopens when someone replies, and whether resolved threads are hidden by default — and a hidden-by-default thread is how a question gets missed. A parent pointer alone has none of those questions, and none of the value.

What a comment body is

How many representations of the text are stored, and which one is the truth?

  • Plain text only. Nothing to sanitise beyond escaping on render, searchable with ordinary tools, and portable anywhere the text needs to go — a notification, an export, a report. It cannot express emphasis, structure, or a mention as anything other than characters, which means addressing someone can only ever be guessed at by parsing.
  • A structured document only. Formatting and, crucially, mentions as nodes with resolved identifiers rather than as text that looks like a name. Everything that wants text now has to flatten the document, and every consumer flattens it slightly differently.
  • Both, stored together. Each reader takes the representation it wants. The cost is that two representations of one thing can disagree, and they will: an edit that updates one and not the other, an import that supplies only one. A project storing both must name which is authoritative, derive the other from it rather than accepting both from the client, and say what a reader does when they conflict.

Independently of that choice: any structured body is validated against a closed allow-list of permitted constructs, and rejected rather than silently stripped when it fails. Stripping is friendlier and worse — it changes what someone said without telling them. The allow-list is enforced on the server; the editor applying the same list is a convenience, not a control. Size is bounded, and for a structured body that means bounding nesting depth and node count as well as byte length, because a document can be small and still ruinous to walk.

Addressing people

  • A comment can address one or more targets. What may be addressed is a decision: individual people only, or also named sets of people — and if sets, whether addressing a set the author does not belong to is permitted. Restricting it prevents an outsider paging a team they know nothing about; permitting it is what makes the feature useful for asking another department a question.
  • Whether the addressed set is extracted into rows or re-derived from the body on every read is the classic derived-data trade. Rows give an indexed "addressed to me" query, a fan-out source that does not re-parse anything, and reporting. They also introduce a second copy that must be regenerated on every edit and can drift from the body — and the drift is silent, because nothing compares them.

When is an address resolved, and against what?

Resolving at write time — checking that each target exists, is visible to the author, and is permitted — turns a mistake into an error message while the author is still looking at the composer. It also snapshots an authorisation: whatever was true when the comment was written.

Resolving at delivery or read time keeps the check current, at the cost of a comment that was accepted and then quietly addresses nobody.

Most projects check on write. The consequence to record rather than discover: membership changes between writing and delivery, so a set addressed on Monday may deliver to a different set of people on Tuesday, and someone who has since left may still be addressed. Neither behaviour is wrong; the one that is wrong is not having decided.

The picker and the write-time check answer the same question with the same predicate.

"Who may be addressed here" is one definition — in scope, active, actually entitled — and it is evaluated in two places: once to fill the picker, once to accept the write. Two hand-written versions of that predicate will drift, and the drift has one shape: a query that reads only the assignment rows for the picker, and a stricter one that also checks the assignment is live for the write. The result is the picker offering a name the server then rejects, which the reader experiences as a broken feature and the developer as an unreproducible bug. The two paths call one membership predicate, owned by the people-and-groups side, and a test lists what the picker returns and asserts the validator accepts every entry.

Notification, and what a successful write promises

Does the write succeed when the notification fails?

  • Notify within the write. A stored comment implies the people addressed were told. It also means the notification path's availability is the commenting path's availability, and a downstream outage stops people talking to each other.
  • Notify after the write, failures logged and swallowed. Commenting always works. The guarantee goes: a success response now means stored, not delivered, and nothing on the screen distinguishes the two — the author believes they have asked someone for something.
  • Record the intent to notify in the same write, deliver separately. Keeps both, at the cost of another moving part and delivery that is eventually rather than immediately true.

Whichever is chosen, say so where a reader will see it, because the whole point of addressing someone is the expectation that they find out. The delivery side of this — channels, preferences, retries and what a recipient is allowed to be told — belongs to Notifications, not here.

Editing, removal, and the difference between them

  • An edit that leaves no trace is indistinguishable from the reader having misread — and in a discussion others have replied to, it can make the replies nonsensical. At minimum a reader can tell that a comment has been edited since it was written.

What does removal mean?

  • Hard delete. The row is gone. Honest, simple, satisfies an erasure request completely, and leaves replies dangling and the conversation unreadable. Nothing can answer "what did it say?", including a legitimate investigation.
  • Soft removal with a placeholder. The comment is retained, readers see that something was removed and by whom, and the shape of the conversation survives. The content still exists, which is the point and also the problem — see Legal Context.
  • Soft removal with content readable under a separate grant. Adds moderation review and support to the above. It also creates a class of reader who can read text that everyone else believes is gone, which needs to be a deliberate, named grant rather than a side effect of being an administrator — and a grant that is named but seeded into the general administrator set by default is a side effect with a label. It is granted on its own, to a role that exists for it.

Three questions follow whichever is chosen, and each is answered explicitly or answered by accident: what happens to the replies of a removed comment; what happens to comments when their target record is removed; and whether removal is available to the author, to a moderator, or to both under different grants — because "the author withdrew this" and "a moderator took this down" are different facts and readers need to tell them apart.

Is there a history of what a comment said, and who may read it?

  • None. Current text only, plus an edited marker. Cheapest, and there is no way to answer what was changed.
  • Full append-only history, one entry per state. Answers it completely: who changed what, when, and what it said before. It also means every removal is only a removal of the current text, and the history is now the most sensitive store in the feature.

The history overlaps a system-wide change trail if the project has one, and two histories of the same object that can disagree is worse than either alone — see Auditing, Reporting & Measurement, and Versioning for the general shape of immutable version history.

A comment written by a business process is a lifecycle decision, not a volume footnote.

A decline reason, a rejection note, a "returned for correction" remark: a process that records its outcome as a comment has written a business record into the discussion stream, and the stream's ordinary rules then apply to it — the row carries the acting user as its author, and the author may edit or remove their own comments. Nothing downstream distinguishes it, so the person who declined a record can afterwards rewrite or remove the recorded reason under an ordinary comment grant.

Decide, per kind of system-written comment, which of these it is:

  • A normal comment the author owns. Editable and removable like any other. Acceptable only if the reason is not the record — something else holds the decision and the comment is a courtesy.
  • Immutable. No edit or removal by anyone, including the author and a moderator; correction is a new comment. This is the default for a comment that is the record of a decision.
  • Editable by a moderator only. The author cannot touch it; a named grant can, and the history shows that it did.

Say which, and enforce it on write — in the update and removal paths, keyed on the kind of comment — not by hiding the affordance in the composer. A vocabulary of comment kinds exists to make this enforceable; a project with one kind of comment has no such rule to write, and says so. A process that writes a comment directly rather than through the ordinary write path also skips whatever that path checks — rate limit, target existence, body validation — and each skipped check is either deliberate and named or a gap.

Who may see what

  • The governing rule is inherited, not local: a reader who cannot see the record cannot see its discussion, under any grant. Comment-specific grants can only ever narrow that, never widen it. Anything else makes the comment stream a way to read a record sideways.
  • Whether writing needs its own grant separate from reading is a project decision with a simple test: are there readers who must not contribute? Read-only auditors, external reviewers and suspended accounts are the usual reasons.
  • The permission model that resolves all of this is Permission Model; the people and named sets that comments address are Users & Groups.

Does one record have one audience, or several?

A stream where every comment is visible to everyone who can see the record is the simplest thing that works, and it fails the moment a record is visible to more than one kind of reader — an internal note next to something a customer, a supplier or an external reviewer can read.

If a project needs more than one audience, the audience is a property of the comment, enforced inside the query that lists them. Filtering after the read, or in the client, produces the single worst failure this capability has: an internal remark rendered to the wrong reader, or counted in a total that tells them it exists. And the composer must make the current audience unmistakable before the text is written, not after — the mistake happens at typing time.

Finding a comment

What is the scope of search: one record, or everything the reader may see?

Within one record is a filter on an already-narrow, already-authorised set. It is cheap even when implemented badly, and it answers the question people actually ask most often.

Across every record the reader may see is a different feature wearing the same word. Results must be filtered by what the reader is allowed to see, which is a per-record authorisation decision, which does not paginate cheaply — the honest implementations either pre-compute the readable set or accept that page sizes are approximate. It also needs a real text index rather than a substring match, and it exposes comment text through a surface that is no longer anchored to a record the reader opened deliberately.

Scoping search to one record is a complete and defensible answer. Saying nothing is not, because the substring match written for one record will be pointed at the whole table by the first person who wants cross-record search.

UI/UX Design ​

The stream is one reusable component wherever it appears, not a per-screen build. Points that matter wherever it is used:

  • Attribution is on every comment, with a stable identity underneath the displayed name, so a renamed or deactivated author still resolves to somebody rather than to a blank.
  • Order is a decision the component makes once, not per screen. Oldest-first reads as a transcript and puts the newest content furthest from the composer; newest-first surfaces the latest and makes replies read backwards. Pick one and apply it everywhere, because a user who learns one order on one screen will misread the other.
  • A removed comment leaves a visible placeholder where soft removal is used, or the surrounding replies stop making sense.
  • The edited marker is derived, not stored. Where a history exists, "edited" is the version number being greater than the first; where it does not, it is the last-change time differing from the creation time. A separate boolean is a third copy of a fact two columns already carry.
  • Timestamps carry a machine-readable value alongside whatever human phrasing is shown. Relative phrasing suits a live discussion; absolute phrasing suits a record being reviewed months later, and a project that needs both shows one and reveals the other.
  • An error state is never rendered as an empty one. "No one has commented" and "we could not load the discussion" are different answers, and a reader who confuses them concludes wrongly about the record. Where the viewer cannot read the host record, the section is absent rather than empty.
  • The composer states the audience and the addressing affordance before text is typed, and addressing offers only targets the author is permitted to address — a picker that offers what the server will reject is a bug report waiting to be filed.
  • Accessibility: the stream is marked up as a list, the composer is a labelled text region with its constraints announced rather than only rendered, and an addressing picker is an accessible combobox — arrow-key navigable, with the active option announced.

Internationalization & Localization ​

  • Comment content is never translated. It is what a person wrote, and a translated quotation is no longer a quotation. A project offering on-demand translation presents it as an addition beside the original, never as a replacement, and never stores it as if it were the comment.
  • Everything the system says around the content — placeholders, the removed-comment notice, validation messages, the addressing picker's empty state, action labels — resolves through the project's translation mechanism and is seeded for every supported locale; see Translations.
  • Timestamps and any relative phrasing are formatted for the reader's locale and time zone, not the author's.
  • Length limits are counted in a unit that behaves the same for every script. A limit that counts bytes silently gives some languages a shorter comment than others.

Non-Functional Requirements ​

  • Every list and search read is paginated. There is no unbounded read of a comment stream, ever — the busiest record in any system is the one someone will open on the worst day.
  • Ordering is deterministic. Time alone is not a total order, so the sort carries a tiebreaker on the identifier, or a page can drop or repeat a comment while someone reads it.
  • The keyset is over a column that never changes after insert — creation time or a monotonic identifier — plus the tiebreaker. A total order is not enough: a cursor keyed on the last-change time is still deterministic and still breaks, because an edit or a soft removal moves the row past the cursor of everyone mid-way through the stream — shown twice to one reader, never to another. "Recent activity first" is a presentation choice layered on top of a stable cursor, not the cursor's key.
  • Where the system scopes its data, comments carry the scope and every read is isolated by it, including the addressing picker — which is otherwise a directory listing with a search box.
  • Writes are rate-limited per author. A comment endpoint is an authenticated, cheap, unbounded-text write, which is the shape of every accidental and deliberate flood.
  • Retention is decided rather than defaulted, and the answer interacts with the erasure questions in Legal Context rather than being independent of them.

Performance ​

  • The dominant read is "the discussion on one record", so it is served by an index on the target reference plus the sort key, and it must not require reading anything about other records.
  • A threaded read is the N+1 trap of this capability: fetching a page of top-level comments and then querying each one's replies, addressed targets and author separately. The shape that avoids it is one query per level, not per row — the page of parents, then their replies, then the derived sets for both.
  • Assembling the tree can happen in the query or in the application. In the query it is one round trip and a harder statement to change; in the application it is simple code and more round trips. Either is defensible at the sizes this feature normally reaches.
  • Substring matching over comment text cannot use an ordinary index. That is acceptable while search is confined to one record and is the first thing to break when it is not.
  • Comment volume is proportional to human effort, not to machine throughput, so this is rarely the system's largest table. The exception is a project that also lets automated processes comment, which changes the growth curve and should be recognised as such when it is introduced — and which raises the lifecycle question under Functional Requirements before it raises a volume one.

Transactional Operations ​

A comment write is rarely one row. It is the comment, plus the derived set of addressed targets, plus a history entry where one exists — and these must not be able to exist without each other. A comment whose addressed targets failed to write addresses nobody while claiming to; a history missing its first entry cannot answer what the comment originally said.

  • The comment and everything derived from it commit together. Partial success here is not a degraded result, it is a wrong one.
  • Notification is outside that boundary under most of the options above; which one a project chose is stated under Functional Requirements and is the reason a success response may or may not promise delivery.
  • Concurrent edits of one comment need a rule. Last-write-wins is the common answer and is acceptable for text one person owns; where a history exists, the version identifier is unique per comment so two concurrent edits cannot both claim to be the same version. The uniqueness constraint is the backstop, not the design: the loser receives a conflict response, which means the write checks the expected version inside the transaction, or the constraint violation is caught and mapped — an unmapped unique-key error surfacing as an internal error is a correct rejection reported as a fault. Where moderation and authorship can act at once — an author editing while a moderator removes — the removal wins, and the edit is rejected rather than silently applied to a removed comment.

Diagrams & Models ​

Writing a comment that addresses people, with the boundary that decides what the response promises:

sequenceDiagram
    autonumber
    actor A as Author
    participant S as Comment service
    participant D as Store
    participant N as Notification channel

    A->>S: Submit a comment on a target
    S->>S: Validate the body against the allow-list and the size bounds
    S->>D: Confirm the target exists and the author may read it
    alt target missing or not readable
        S-->>A: Rejected — the target is not available
    end
    opt replying
        S->>D: Load the comment being replied to
        alt wrong target or too deep
            S-->>A: Rejected — the reply is not permitted
        end
    end
    S->>S: Extract the addressed targets from the body
    S->>D: Resolve each target and check the author may address it
    alt unresolved or not permitted
        S-->>A: Rejected — naming the offending target
    end

    rect rgb(245,245,245)
        Note over S,D: One transaction — the comment,<br/>its addressed targets, its first history entry
        S->>D: Write all three
    end

    S->>N: Hand off delivery to the addressed people
    Note over S,N: Outside the transaction in most designs,<br/>so a success below means stored, not delivered
    S-->>A: Accepted

API Analysis (API-A) ​

No separate API-A page. Unlike a capability whose surface carries decisions its feature page does not, the read and write surface here is ordinary create / read / update / remove over one collection, and every choice worth making about it is already stated above: the attachment model fixes the filter, the conversation shape fixes what a page is a page of, the audience decision fixes what the list query may return, and the search scope decision fixes whether a second read surface exists at all. A project instantiating this concept writes its own API-A against its conventions; a second, hypothetical one here would restate the decisions and then have to be kept true twice.

Two surface-level points do not follow from anything above and belong somewhere:

  • The addressing picker is an endpoint, and it is a directory search. It returns who exists and what they are called, to anyone who may write a comment. Scope it, bound it, and treat its results as an enumeration surface rather than as a convenience — see Cybersecurity Considerations.
  • Removal is not an update. Modelling it as "edit this field" makes moderation indistinguishable from authorship in every log, permission check and history entry that follows.
  • A malformed cursor is rejected, not silently ignored. Treating an unparseable cursor as "first page" turns a client bug into a quiet reset — the reader sees the stream start over and nobody sees an error. It is a validation failure on the request, answered as one.

Domain Model ​

  • The comment — the target reference, the author, the body in whichever representations the project stores, the creation and last-change times, the lifecycle marker if removal is soft, the scope where the system is scoped, and the conversation linkage (a parent reference, a thread reference, or both). Plus whatever identifier and audit columns the project's conventions give every entity.
  • The addressed set — optional, derived from the body, one row per resolved target carrying the kind of target and its identifier. Present only where addressing is extracted rather than re-parsed.
  • The history — optional, append-only, one entry per state of the comment, carrying what the body was, what act produced the entry, who performed it and when. Exactly one entry per (comment, version); never updated, never deleted.
  • The thread — optional, present only where a thread is first-class, carrying whatever state the conversation has that no individual message owns.
  • Vocabularies, each declared once and closed: the kinds of record that can be commented on, the kinds of thing that can be addressed, the acts that produce a history entry, the audiences a comment can belong to, and — where a business process writes comments — the kinds of comment, with which kinds are user-owned and which are records under the lifecycle rule stated above.

Logging ​

.info

  • a comment was written, edited or removed (identifier, target reference, actor, and — for removal — whether the actor was the author or a moderator, as a recorded field on the log line and on the history entry, not something a reader derives later by comparing the removing actor with the creating one).
  • a search was executed (target, actor, and the length of the term, not the term).

.warn / .error

  • a write failed, an addressed target could not be resolved, a stored body failed to parse against the current allow-list, or a notification hand-off failed. The last one is, under the fail-open option, the only trace that anyone was not told.

Comment text is never logged, in whole or in part, and neither is a search term. Logs are read by people and systems that hold no grant over the record the comment is attached to, and a log line is the easiest way to move freeform personal data outside every control this feature has.

Monitoring ​

Failed writes and failed notification hand-offs deserve an alert; under the fail-open option the second is the only signal that people are not being told they were addressed. Rate-limit rejections are worth watching as a product signal rather than an incident — a rising rate usually means the limit is wrong rather than that anyone is abusing anything.

Caching ​

None, ordinarily. A comment stream is the part of a record most likely to have changed since the reader last looked, and a stale discussion is worse than a slow one: it hides the reply that was just written. The addressing picker is the one candidate for caching, because it reads a directory that changes far more slowly than the discussion does — and even there the cache must be scoped, or it becomes a way to see who exists outside the reader's scope.

Indexing ​

Every filter the read surface offers is backed by an index: the target reference plus the sort key for the stream itself, the parent or thread reference for assembling replies, and the addressed target for "addressed to me". Where removal is soft and most reads exclude removed comments, those indexes are partial over the rows that remain — but only where the retained rows are genuinely excluded, because a threaded view that renders placeholders needs them present. A history table is indexed by the comment it belongs to, with the version uniqueness enforced there rather than in application code.


Comments are freeform text written by people about business records and, often, about other people. That combination means the legal questions here are less predictable than for structured data, where the fields are known in advance. What binds a given project depends on jurisdiction, sector and contract, so these are questions to settle:

  • Is the discussion part of the record for retention or disclosure purposes? A regime that requires a decision to be retained may reach the reasoning behind it, and comments are frequently where the reasoning is. If they are in scope, they inherit that regime's retention floor rather than the project's convenience.
  • Does an erasure obligation reach comment content — and how far? The body is the obvious part. The history, the addressed-target rows, the notifications already delivered and any export or backup already taken are the parts that get forgotten. A soft-removal design keeps the content by design, so erasure and removal are two different operations and a project that has only built one should know which.
  • Does an author have the right to withdraw their own words? And does that right survive other people having replied to them, or been notified of them?
  • Can a comment ever reach someone outside the organisation — through a shared view, an export, a printed record, a notification to an external address? If it can, is that visible to the author before they write, and does anything constrain what may be written where?
  • Is comment content processed by anything beyond storage and display — translation, classification, summarisation, model training? Each is a processing purpose, and a purpose usually needs a basis and a notice.
  • Does an employment or works-council regime constrain what may be derived from who commented? Volume, response time and tone are all computable from this data and are all measurements of people.
  • Where is the content stored and replicated, and does any residency commitment cover it? Freeform text is the category least likely to have been considered when that commitment was made.

Cybersecurity Considerations ​

Comment bodies are untrusted input from authenticated users, which is the category most often under-defended because the author is a known colleague rather than an anonymous stranger. The text is treated as data everywhere, never as code:

  • A structured body is validated against a closed allow-list of permitted constructs, on the server, on every write, and rejected rather than stripped when it fails. The allow-list is the primary control, so a gap in it is the primary vulnerability. Size, nesting depth and node count are bounded independently of the allow-list, because a permitted construct repeated enough times is its own attack.
  • A plain body is stored and rendered as text, escaped at render, never interpreted as markup.
  • Addressing attributes are validated as identifiers, not accepted as free text, so an address cannot smuggle markup or a lookup through its own fields.
  • Rendering happens in one place. A second renderer — an export, a notification template, an administrative view — is a second place the allow-list has to be honoured, and the one that will not be.

The addressing picker is an enumeration surface. It answers "who exists and what are they called" to anybody who may write a comment, which is a broader audience than any directory screen usually has. It is scoped, bounded in the number of results, rate-limited, and returns the minimum that makes the picker usable — a picker that returns contact details returns them to everyone who can type into a composer.

Notification carries content across the permission boundary. A notification containing the comment body delivers text about a record to a channel that never re-checks the recipient's access to that record — and being addressed in a comment is not the same as being permitted to read what it is attached to. A project either verifies access at delivery time or sends a pointer rather than the content. A preview or excerpt is content for this purpose — the first few hundred characters of a body carry the same risk as the whole of it and are stored by the notification side exactly as the whole would be. Sending the body because it makes a nicer email is how a record leaks to someone who was never entitled to it.

Data Privacy. This is the least structured personal-data store in most systems: it collects information nobody modelled, about people who are not the subject of the record, in a field with no schema. Access is bounded by access to the target record, scope isolation applies on every path including search and addressing, and text is kept out of logs. Where removal is soft and content is readable under a separate grant, that grant is named, narrow and its use is itself observable — otherwise "deleted" means one thing to users and another to the system.


Risk Assessment ​

Business Risks ​

Direct exposure is small — commenting is a collaboration aid, and a system whose comments are unavailable continues to do its work. The indirect exposure is larger and slower: once discussion lands here it becomes the record of why, so losing it, hiding it behind a permission change, or ageing it out on a retention rule chosen for storage reasons destroys institutional memory that was never written anywhere else.

The sharpest business risk is a single incident rather than an aggregate: a remark reaching a reader it was not meant for. Internal candour about a customer, a supplier or a colleague is the normal content of an internal note, and one misrouted note costs far more than the feature is worth. That is the whole reason the audience decision above is enforced in the query rather than in rendering.

Technical Risks ​

  • A gap in the body allow-list is the primary security surface, and it is exploited by a colleague's browser rather than by an attacker's. Mitigations: server-side enforcement, rejection rather than stripping, independent size and depth bounds, and a single rendering path.
  • Dangling target references where the attachment is polymorphic. Readers tolerate a comment whose target no longer resolves; a surface that assumes resolution fails on exactly the records someone is trying to explain.
  • The derived addressed set drifting from the body, because the body was edited through a path that did not regenerate it. The drift is silent — nothing compares the two — and it presents as people mysteriously not being told.
  • Search that cannot use an index, correct and cheap at one record's scale, quietly repointed at the whole table the first time someone wants cross-record search.
  • The N+1 threaded read, which passes every test written against a record with four comments.
  • Coupling to notification, where the failure mode depends entirely on the decision recorded above and, under fail-open, is invisible to the author.

Auditing, Reporting & Measurement ​

Two histories of the same object is the trap specific to this capability. A comment that keeps its own version history, in a system that also keeps a general change trail, has two records of the same edit that can disagree — and a reader has no way to know which to believe. There are four honest resolutions: the comment keeps its own history and is enumerated as an exemption from the general trail, with the exemption written down and countable, as Audit Log requires; or it keeps no history of its own and the general trail is the only answer; or both exist and one is declared authoritative, which is only credible if the other is derived from it rather than written independently; or the two are partitioned by question — the version store holds what was said, under its own entitlement, and the trail holds who did what to which record with a version pointer and no body, with the boundary written down and the version number as the join key. What is dishonest is two stores answering the same question independently, which is the outcome nobody chooses and many projects have.

Moderation is the part that most needs recording, and it is the part most often folded into an ordinary update. "The author withdrew this" and "someone else took this down" are different facts with different consequences, and a design that records only "removed, by actor X" makes them indistinguishable the moment the moderator is also an author. Recording the reason for a moderation removal is a further decision — it is the only defence a moderator has, and it is another piece of freeform text inheriting every question in Legal Context.

Whether reading is recorded is a separate decision from whether writing is. Comments frequently contain the most sensitive freeform text in the system, so who read a discussion can matter more than who wrote it. Most projects record nothing and never say so, leaving a reader unable to tell a decision from an oversight.

What is measurable, and the caution attached to it. Comment counts per record, the share of records carrying any discussion at all, addressing volume, time to first reply, and — only where threads are first-class — the age of unresolved threads. The first two are the honest measures of whether the capability is used. The rest are measurements of people as much as of the system, and per-person versions of them are the ones Legal Context asks about before they are built, not after.

Reporting is browsing for most projects, and that is a complete answer. Anything aggregate is a separate build with the usual choice — against the live data, where it competes with the read path, or against a copy, which costs a copy and a lag. Saying nothing is what is not complete, because the first ad-hoc query someone writes against comment text will be written against production.

Signals worth watching, each of which only this feature can produce: the rate of rejected bodies, which distinguishes an allow-list that is too narrow from an attack; the rate of failed notification hand-offs, which under the fail-open option is the only evidence that people were not told; the share of comments that are edited or removed shortly after being written, which usually indicates a composer problem rather than a discipline problem; and the use of any grant that reads removed content, which is the highest-privilege read this feature has.

Known instances ​

None recorded in this copy. Add this project's instance here as it adopts the concept — one line naming the feature that implements it. This is the only place a concept may name a project, and the template ships it empty so a fork inherits no other project's entries.