From the source of truth

RFC 0003 - Delivery semantics

A static snapshot from the Sagüin repository.

, so a key that does holds no channel's value - and it is not broadcast either, since a publish there is refused too |\n| the key is empty, or longer than `max_topic_length` or deeper than `max_topic_levels` | refused, the first `0x83` and the others `0x90` |\n| no Response Topic | refused `0x83`. A read with nowhere to send its answer has no effect at all, and accepting it would tell the client it had worked |\n| QoS 0 | dropped. There is no `PUBACK` to carry any of the above |\n\nThe two codes are split by what is wrong rather than by where the check\nhappens, because the code is what `saguin_publish_refused_total` carries\n(RFC 0005): `0x83` says the request cannot be acted on, `0x90` says the key\nis not a topic this broker has a value for. An operator watching one number\ncan then tell a fleet mistyping keys from a client that forgot its reply\npath.\n\n**Every reply carries Correlation Data**, whether or not the caller sent\nany: the caller's when it did, and otherwise the key it asked about. A\nclient with several reads outstanding can tell the answers apart without\nkeeping a table, and one that wants its own bookkeeping still has it.\n\n**The reply goes to the asking session and to nobody else**, exactly as a\nseek's does and for the same reason: it is written to the connection that\nasked, needs no subscription, and a third party on the same Response\nTopic sees nothing.\n\nThe consequence to know before trying it: `mosquitto_pub` publishes and\nexits, so it asks and is gone before it can show you the answer - `-d`\nprints the reply arriving, and anything more than that wants a client that\nstays connected.\n\n**A point read never makes the caller a subscriber**, which is the whole\nreason it exists. Subscribing to ask returns silence for an absent key -\nindistinguishable from a slow one - and enrols the caller in every later\nupdate to a topic it wanted once.\n\nIt reaches `latest` channels only. An `append` channel is a log rather than\na set of current values, and a queue holds unresolved work that exactly one\nconsumer may take.\n\n### Saying which schema deserializes a payload\n\nA publisher can already say *how* its payload is serialized: Content Type\nand Payload Format Indicator are MQTT 5's own fields and Sagüin carries\nboth. What it cannot say with them is *which schema* -\n`application/x-protobuf` does not tell a WeatherReading from a WaterLevel,\nand a consumer holding the bytes has no way to find out.\n\n**This needs nothing from the broker.** A `latest` channel is already a\nkey-value store with delete, readable one key at a time by the point read\nabove, so a schema registry is a channel and a convention:\n\n```yaml\nchannels:\n schemas:\n type: latest\n filter: schemas/#\n```\n\n- **Register** by publishing the schema text to a topic in that channel.\n- **Retire** by publishing a zero-length payload, which is how this channel\n type deletes.\n- **Produce** with Content Type saying the serialization format and a\n User Property named **`schema`** carrying the schema's topic.\n- **Consume** by reading that property, point-reading the topic it names,\n and caching the answer until the pointer changes.\n\n**The property is called `schema`**, and Sagüin recommends the name rather\nthan leaving it to each deployment for the reason a convention exists at\nall: a consumer written against one name works against any broker that\nfollows this section, and a consumer written against a name somebody chose\nlocally works nowhere else. Nothing in the broker reads it - it is an\nordinary User Property and Sagüin never touches it - so an operator with a\nreason may use another name, and pays for it in portability.\n\n**The pointer is a whole topic rather than a bare id**, and that is what\nmakes the convention work without an allocator. A bare `weather-v1` has two\nholes: two publishers in different domains choose the same name and the\nsecond write silently replaces the first, and nothing in the message says\nwhich channel or prefix to look under. A topic answers both - the consumer\nreads exactly the string it was given.\n\n**Collisions are refused rather than discouraged.** One ACL rule confines\neach publisher to its own prefix, so two domains choosing the same name\ncannot reach each other's key:\n\n```yaml\nroles:\n device:\n - channel: schemas\n filter: \"schemas/%u/#\" # acme may write only under schemas/acme/\n allow: [write, read, delete]\n reader:\n - channel: schemas\n allow: [read] # read any schema, write none\n```\n\n**The registry can live anywhere.** `filter:` is compared against the\nwhole topic, so a registry at `schemas/#` in a channel named `schemas` and\none nested at `iot/+/schemas/+` are written the same way - the second's\nrule is `filter: \"iot/%u/schemas/+\"`, and nothing about the mechanism\nchanges.\n\n**A consumer follows a pointer only where it lands inside the registry's\nown filter.** The ACL governs who may *write* a topic and never who may\n*name* one, so a publisher may point at anything - including a `latest`\nchannel holding device state. The check is one comparison and the filter is\nreadable from `saguin_channel_info`.\n\n**`saguin-` is reserved on a publish**, so `saguin-schema` - the name that\nlooks most official - is the one that will not survive: it is dropped like\nevery other reserved name a client sets, silently, because a publisher\ncannot be allowed to forge the broker's own properties.\n\n**What Sagüin does not do**, said plainly because a registry usually does\nit: it does not parse a registered schema, check that it is valid, or\nenforce that a new version is compatible with the last. The payload is\nopaque bytes here as everywhere else. A deployment that must refuse an\nincompatible evolution still needs something that understands the format.\n\n## When retention removes\n\nBoth retention rules are applied by a sweep on the broker's clock rather\nthan by the write that violates them, and neither interval is\nconfigurable.\n\n**A channel may therefore hold slightly more than it says, briefly.** Over\nits `retention_bytes` until the next sweep, and past its\n`retention_period` by about a tenth of that period. That is the difference\nbetween retention and `max_bytes`: `max_bytes` is a refusal a producer is\ntold about with `0x97`, so it holds exactly, and retention is a\nhousekeeping target that removes behind a producer and never says\nanything. A bound nobody is told about does not have to be exact; it has\nto be reached.\n\nThe alternative was trimming on the publish that goes over, which is exact\nand is paid on every publish to a channel at its size: a permanent tax on\nthe write path to make a number exact that nothing reports. Kafka makes\nthe same trade with `log.retention.check.interval.ms`, and for the same\nreason.\n\n**Age needs a clock whatever size does**, because a channel that has\nstopped receiving writes still ages, and no publish will ever arrive to\nnotice.\n\nThe two run on separate clocks because they decide different things. The\nsize sweep decides how far above its number a channel may sit, so it is\nfrequent and fixed. The age sweep decides only how far past its deadline\na record may survive, so it is a tenth of the shortest deadline any channel\nconfigured - a month's retention is served perfectly well by a sweep every\nthree days. **A queue's `job_expires_after` counts as one of those\ndeadlines**, since it is the same sweep that enforces it, and the whole is\nfloored at one second, which is also what a configuration carrying no\ndeadline at all gets. Both sweeps run once at startup too, so a broker that\nwas down while its records aged does not wait for a tick to notice.\n\nNone of this is exposed. What an operator would be choosing is how far\npast their own policy they are willing to sit, and there is no answer to\nthat except \"as little as it costs\" - which is what the broker already\npicks. It states both intervals at startup.\n\n## `queue`\n\n### States\n\n```\n ┌──────────────┐\n publish ────────────▶│ AVAILABLE │◀──────────┐\n └──────┬───────┘ │\n PUBLISH │ │\n ▼ │\n ┌──────────────┐ │\n │ DELIVERING │───────────┤ disconnect\n └──────┬───────┘ │ before PUBACK\n │ │ never sent\n PUBACK │ │\n ▼ │\n ┌──────────────┐ return │\n │ LEASED │───────────┤ timeout\n └──────┬───────┘ │ disconnect\n │ │\n ack │ │ (attempts remain)\n ▼ │\n ┌──────────────┐ │\n │ RESOLVED │◀─────────┘ attempts exhausted\n └──────────────┘ ⇒ dead-letter\n```\n\n`DELIVERING` and `LEASED` are separate on purpose: a record's visibility\ndeadline must never overlap its transport in-flight window. While\n`DELIVERING`, the record occupies a slot in the worker's MQTT in-flight\nwindow and no visibility deadline is running. The deadline starts at\n`PUBACK`, when that slot is released. A record can therefore never be\nsimultaneously held unacknowledged at the transport layer and expired at\nthe application layer, which is the condition that would otherwise leak a\nflow-control slot per timeout until the worker silently stopped receiving\nwork.\n\n**`DELIVERING` has no deadline, so it needs the other exit: a record whose\n`PUBLISH` was never put on the wire is returned to `AVAILABLE`, with its\nattempt unspent.** Waiting for the `PUBACK` assumes one can arrive, and for\na record the connection never carried none ever can - so without this the\nrecord stays `DELIVERING` until the worker's session ends, held by a worker\nthat has never seen it: not redelivered, not dead-lettered, and offered to\nnobody else.\n\n**It is not a deadline on `DELIVERING`** - \"no `PUBACK` yet\" is the\nordinary state of a job in progress, so any clock short enough to rescue\na stranded record takes a live worker's job away (invariant 7). Whether\nthe packet was sent is a question with an answer, so the broker asks it\nrather than timing it, and a worker mid-job is never touched.\n\n### Delivery\n\n**A worker subscribes to `$saguin/queue/\u003cchannel>` at QoS 1**, and nothing\nelse consumes a queue. **A subscription asking for No Local is refused\n`0x8F`**, naming the filter to change; the connection stands. The reason\nis the one MQTT gives for making the same flag a Protocol Error on a\nshared subscription: echo-suppression and work distribution cannot both be\nhonoured. A job withheld from the one worker whose client id published it\nis never delivered - so it is never acknowledged, never times out, never\ndead-letters, and nothing reports a job that has stopped moving. A single\nworker publishing its own work is all it takes. The broker keeps its own\nindex of who is subscribed to which queue and selects one of them -\nround-robin over the workers that are live, hold no job from this queue,\nhave room in their in-flight window, and are still allowed the channel by\nthe `acl_file` - then sends a `PUBLISH` at QoS 1 carrying:\n\n| Property | Value |\n|---|---|\n| Response Topic | `$saguin/queue/\u003cchannel>/response` |\n| Correlation Data | the Delivery ID for this attempt |\n| User Property `saguin-id` | the Message ID |\n| User Property `saguin-offset` | the record's position in the queue |\n| User Property `saguin-timestamp` | when the broker received it, Unix milliseconds |\n| User Property `saguin-expires` | when the publisher's expiry runs out, Unix milliseconds; only where one was set |\n| User Property `saguin-attempt` | this attempt's number, from 1 |\n\nThe four `saguin-` properties before the attempt are on every channel\ndelivery, `saguin-expires` only where an expiry was set; Response Topic,\nCorrelation Data and `saguin-attempt` are the queue's own. On a queue\n`saguin-expires` is the one a worker is expected to act on: nothing removes\na job on the publisher's clock, so a job whose moment has passed is still\noffered, and whether stale work is worth doing is the worker's decision\nrather than the broker's.\n\n**The topic and the payload are the publisher's, unchanged, and so are its\nown User Properties.** A delivery is the original message with the\nproperties above added to it - there is no envelope, nothing to unwrap, and\nnothing rewritten. `orders/resize/thumbnails/42` is delivered on\n`orders/resize/thumbnails/42`, and a header the publisher set arrives beside\n`saguin-id` rather than inside anything. Only the `saguin-` prefix is\nreserved (see \"From publish to record\").\n\n**The topic under a queue channel does not select a worker, and this is the\none place a habit from ordinary MQTT misleads.** On an `append` or `latest`\nchannel a subscriber's filter chooses what it receives. A queue admits\nexactly one subscription form, so every worker draws from the whole channel\nwhatever the topic says: it is description, not address. It survives storage\nand redelivery, and it carries into the dead-letter channel - a record\npublished to `orders/resize/thumbnails/42` is dead-lettered to\n`orders/__dlq/resize/thumbnails/42`, which is what lets an operator see\nwhat failed.\n\nWork that must go to different pools of workers is therefore **two channels,\nnot two filters**, and the broker says so: a narrower filter is refused with\n`0x8F` rather than granted and quietly given a share of everything.\n\n**A worker holds one job from each queue at a time.** A job is held from\nthe moment it is offered until the worker acknowledges or returns it, its\nvisibility timeout runs out, or the worker disconnects, and until then\nthat worker is offered nothing more from that queue. A worker consuming\ntwo queues holds one job from each, and the two do not wait on each other.\n\n**Receive Maximum does not change this.** It bounds what the connection\nhas in flight, a job included, and a worker whose window is full is passed\nover as one holding a job is. More concurrency is more workers, and Sagüin\nadds no configuration for it.\n\n**An answer offers that worker its next job at once**, and so does the\nacknowledgement that reopens a full window - a client may answer a job\nbefore it acknowledges the packet, as Paho does. New work, and a worker\nthat has just subscribed, are offered on a fixed 200ms tick, so a first job\ncan wait that long. With `examples/support/worker` doing no work per job,\none worker drained a 300-job backlog in 53 to 213ms over six runs at\nReceive Maximum 1, 2 and unset, the longest being those whose first offer\nwaited for the tick. A worker is bounded by its own processing, not by the\nbroker.\n\nThe tick is not configurable, for the reason the sweep intervals are not -\nan operator has no way to reason about what to set it to, and the honest\nanswer is \"as little as it costs\".\n\nThe **Delivery ID** is opaque bytes, at most 32 of them. Clients echo it\nexactly and never parse it. It is unique to the attempt: a redelivery of\nthe same record carries a different one.\n\n### Acknowledgement and return\n\nThe worker publishes to the Response Topic it was given, echoing the\nCorrelation Data it was given, with a payload of exactly:\n\n| Payload | Meaning |\n|---|---|\n| `ack` | the work succeeded; resolve the record |\n| `return` | the work failed; make it available again now |\n\nThat is the entire protocol. It is an ordinary `PUBLISH`, and this is its\nshape - though as a command it cannot work on its own: `mosquitto_pub` is\na second session, so rule 3 below has the broker acknowledge it and\nignore it; the job comes back to its worker as attempt 2 when the\nvisibility timeout passes.\n\n```sh\nmosquitto_pub -V 5 -q 1 -t '$saguin/queue/jobs/response' -m ack \\\n -D publish correlation-data \"$CORRELATION_DATA\"\n```\n\nIn a worker it is the message handler answering on the session the job\narrived on - Paho, which acknowledges once and is not redelivered:\n\n```python\ndef on_message(client, userdata, job):\n props = Properties(PacketTypes.PUBLISH)\n props.CorrelationData = job.properties.CorrelationData\n outcome = \"ack\" if process(job) else \"return\"\n client.publish(job.properties.ResponseTopic, outcome, qos=1,\n properties=props)\n```\n\n`PUBACK` is not acknowledgement. It confirms MQTT moved a packet and\nsays nothing about whether the work succeeded. A worker whose library\nacknowledges on receipt has told Sagüin only that it received the job.\n\nThe broker accepts the response only when **all** of the following hold:\n\n1. Correlation Data is present and decodes to a known delivery;\n2. that delivery is the record's **current** one;\n3. the session publishing the response is the session holding it.\n\nRule 3 is what stops one worker resolving another's work. Rules 1 and 2\nare the fencing check, and it runs on every operation naming a delivery,\nnot only on `ack`.\n\nA response failing any of them is ignored and logged. A payload that is\nneither `ack` nor `return` is ignored and logged. Ignoring is safe: the\nlease expires and the record is redelivered.\n\n**A worker is not told whether its response was applied.** The `PUBACK` it\ngets confirms the broker received the packet, nothing more. If the\nresponse was too late, another worker already holds the record and there\nis nothing useful the first worker could do about it. Adding a reply path\nwould be a second protocol for no decision anyone can act on.\n\n### Attempts, timeout, and redelivery\n\nThe attempt count increments when the worker is known to have received the\nrecord: at `PUBACK`, or when it answers for that delivery, whichever comes\nfirst. A delivery that reaches neither burns nothing - the worker never\nshowed it had the job.\n\nBoth halves are load-bearing. `PUBACK` alone is not enough, because a\nclient library sends it when its receive handler returns, so a worker that\nanswers from inside the handler answers *before* it acknowledges. Counting\nonly at `PUBACK` would leave such a worker's `return` uncounted, and a job\nit keeps handing back would be redelivered for ever instead of being\ndead-lettered. Answering is the stronger evidence of the two: a worker\nthat replies about a delivery demonstrably had it.\n\nA lease ends in one of four ways:\n\n| | Result |\n|---|---|\n| `ack` | RESOLVED |\n| `return` | AVAILABLE, after `retry.backoff` |\n| deadline passes | AVAILABLE, after whatever is left of `retry.backoff` |\n| worker's session ends | AVAILABLE, after whatever is left of `retry.backoff` |\n\n**The last three are one rule and not three**, and the next section is what\nit is: the gap runs from the moment the worker last had the record. A\nworker that answers is measured from its answer, so it waits the whole gap.\nA record taken back by the deadline or by a session ending is measured from\nwhen that worker received it, so however long it was held counts toward the\ngap - and at the default 30s visibility timeout, a gap under five attempts\nat a 2s base is already spent.\n\nIn every case that returns the record, the delivery epoch advances, so the\nDelivery ID just used is dead. A late `ack` from the previous holder\nresolves nothing and changes nothing.\n\nA worker's session ending returns every record it held, immediately. The\nbroker knows which records those are, so this needs no waiting for a\ndeadline.\n\nWhen a record would become AVAILABLE but its attempt count has reached\n`retry.max_attempts`, it is dead-lettered instead.\n\n### Backoff\n\nA returned record is offered again on the next tick, which is right for a\njob that failed on its own and wrong for one failing because something\ndownstream is down: the faster the worker fails, the harder the queue\nhammers whatever is already broken, and `retry.max_attempts` is spent\ninside a second. `retry.backoff` puts a widening gap between the attempts.\n\n| `retry.backoff` | The gap after `n` attempts |\n|---|---|\n| `none` - the default | none; the record is offered on the next tick |\n| `linear` | `backoff_base x n` - 2s, 4s, 6s at a 2s base |\n| `exponential` | `backoff_base x 2^(n-1)` - 2s, 4s, 8s, 16s at a 2s base |\n\n**There is no maximum gap and there is no key for one.**\n`retry.max_attempts` already bounds the total: five attempts at a 2s base\ncome to 30 seconds of waiting on the exponential shape, and a ceiling would\nbe a second bound on a quantity that already has one.\n\n**One rule decides every case: the gap runs from the moment the worker last\nhad the record**, and not from the record's `saguin-timestamp`, which is\nwhen the broker received it.\n\nThree things follow from that one rule, and none of them is a rule of its\nown:\n\n**A record taken back by the visibility timeout waits no further.** The\nworker had it from the moment it was delivered, and the timeout is exactly\nhow long ago that was - so a gap shorter than `visibility_timeout` is\nalready served and the record goes out on the next tick. That is the\nintended answer and not a concession: the timeout is a detection delay, but\nit is also real time in which nothing was retried. Where the computed gap\nis *longer* than the visibility timeout, the difference is what such a\nrecord waits, and at the default 30s timeout no gap under five attempts at\na 2s base reaches it.\n\n**A worker whose session ends is the same case**, and it is the one where\nthe arithmetic can still leave something to wait: the record is measured\nfrom when that worker received it, and a session that ends a second after a\ndelivery has served only a second of the gap.\n\n**The backoff and the visibility timeout never compete.** They govern\nopposite states. While a worker holds a record the timeout is what ends\nthat state, and the backoff is not consulted at all - the record is not a\ncandidate for delivery, so there is nothing to delay. Once the record is\nback in the queue the timeout is over and the backoff is what decides when\nit goes out.\n\n**A record still waiting is skipped, and never waited on.** The offer walks\npast it to the records behind it, so one job nobody can process cannot stall\nthe channel for the length of its own gap.\n\nTwo things about the shape of it. The gap is honoured to within one 200ms\ntick, so the wait is the gap rounded up to the next one; and a record that\nis skipped holds no worker, because it is never offered.\n\n### At a size bound\n\nA queue takes `max_bytes` and never retention, because deleting\nunacknowledged work is eviction of unresolved work and is never permitted\n(invariant 2). What a queue does at its bound is refuse the publish with\n`0x97`, and that makes the bound **flow control rather than a limit**:\nresolution is what frees the space, so a queue at its bound accepts work\nagain as its workers catch up, without an operator doing anything.\n\nThe rule that keeps it from being a deadlock is worth stating on its own,\nbecause a bound that refuses every write has one:\n\n> **A bound never refuses the operation that would relieve it.**\n> Acknowledgement, return, the visibility timeout, **the attempt count**,\n> a consumer's stored position, and the move into the dead-letter channel\n> are not publishes and are never refused for want of capacity.\n\nWithout it a full queue could not dead-letter, so it could not drain, so\nit would stay full - and the dead-letter move is exactly the path that\ntakes bytes out of a queue that nothing is succeeding at. A move that\nfails for any *other* reason still leaves the record as it was, with its\nattempt unspent, as above.\n\n**The attempt count is in that list because it is a precondition.** A\nrecord whose attempt count cannot be written never reaches\n`retry.max_attempts`, so it never becomes eligible for the move that\nwould take it out of the queue: the bound blocks the relief indirectly,\nby blocking the thing that has to happen first. It reads as bookkeeping\nand it is a precondition.\n\n**Every one of those operations is a write, so a bound has to hold room\nfor them.** A provider therefore keeps a **reserve** - one largest\npossible record with its headers, carved out of `max_bytes` and derived\nfrom `max_message_size` plus `max_header_bytes` rather than configured -\nwhich publishes may not touch and the operations above may. A `max_bytes`\nbelow twice the reserve is a startup error. How each provider keeps it,\nand why the reserve is spent rather than lent on sqlite, is RFC 0004\n\"The reserve\".\n\n### Dead-lettering\n\nThe record leaves the queue and appears in `\u003cchannel>__dlq` in one\ntransaction, or neither happens. A move refused for want of capacity\nleaves the record exactly as it was, including its delivery state, so the\nattempt is not consumed by a failure that was the broker's.\n\nIts topic gains one level, `__dlq` (RFC 0002 \"The dead-letter channel\"),\nwhich the move writes directly: a record at `limits.max_topic_levels` is\nkept one level past it in the dead-letter channel, and is read there through\na shallower filter ending in `#` (RFC 0002 \"How deep a topic may be\").\n\nThe dead-lettered record keeps its Message ID and gains headers:\n\n```\nsaguin-dlq-channel the queue it came from\nsaguin-dlq-offset its offset there\nsaguin-dlq-attempts how many attempts were made\nsaguin-dlq-first first delivery time\nsaguin-dlq-last last delivery time\nsaguin-dlq-at dead-letter time\nsaguin-dlq-reason attempts_exhausted | expired\n```\n\nThey are ordinary User Properties, so a stock `mosquitto_sub` on\n`jobs/__dlq/#` reads them without any Sagüin-specific tooling.\n\nA record whose age exceeds `job_expires_after` is dead-lettered without\nfurther delivery, with reason `expired`. It goes whatever its attempt count\nsays, because it is out of time rather than out of attempts - a job that\nexpires before any worker exists is dead-lettered on its first offer with\nno attempt spent.\n\n**`job_expires_after` is the only clock that ends a job.** A publisher's\nown Message Expiry Interval does not, on this channel type or any other:\ntaking work out of a queue is the operator's decision, and this is where\nthey make it. A queue that writes no `job_expires_after` never expires work\nat all.\n\n**What a publisher's expiry does instead is arrive and be read.** The\ndelivery carries `saguin-expires`, the moment the interval runs out, and\nthe worker decides what to do about it - ignore the job and acknowledge, or\ndo it anyway and acknowledge. The broker delivers; the application judges.\nThat is the same division as everywhere else here: the broker enforces\nrules an operator wrote, and does not make an application's decisions for\nit.\n\n**This is a deliberate deviation from MQTT and is worth naming.** The\nspecification deletes a message whose expiry has passed before it was\ndelivered. Sagüin does not, on any channel: a record is what an operator\nasked it to keep, and a client does not get to remove it. The one place the\npublisher's clock still deletes is a retained value on a broadcast topic,\nwhich is MQTT's own store, has no consumer position to damage, and is\ncovered under \"Retained messages\".\n\n**Two places check it, and they are not redundant**: a job is checked as\nit is about to be handed to a worker, and the broker's own sweep expires\njobs that are waiting - a queue whose workers have all gone away is never\nasked for anything, and expiry exists for exactly that queue.\n\nThe sweep leaves alone any job that is out with a worker. Removing one is\nremoving a record in flight to a consumer, and the worker's answer would\nthen arrive for a job that no longer exists. Such a job expires on its next\noffer instead, or when the worker hands it back.\n\nA job with no timestamp never expires. Storage keeps an unset time as zero,\nwhich read as a date is in 1754, so a queue of them would be dead-lettered\nwhole on the first check.\n\nThe dead-letter channel is an `append` channel in every other respect:\nindependent consumers, replay, its own retention.\n\n### Putting dead-lettered work back\n\n\"How do I retry the dead letters once the bug is fixed?\" has an answer and\nit is four lines of any MQTT client, because both halves already exist: a\ndead-letter channel is an ordinary `append` channel, and a queue takes an\nordinary publish.\n\n**Read the dead letters.** Subscribe to the dead-letter channel as to any\n`append` channel. A durable session resumes where it left off; a seek to 0\nreplays everything the channel still holds, which is what to use after\nfixing a bug that dead-lettered work over an afternoon.\n\n**Take the `__dlq` level out of the topic.** A dead letter carries the\njob's own topic with one level inserted, so removing it is the way back.\n**Where it sits depends on the queue's filter**, and this is the part to\nget right rather than assume: a filter ending in `#` takes the level at\nthe `#`, and any other filter takes it on the end.\n\n| Queue filter | A job on | Dead-lettered to |\n|---|---|---|\n| `jobs/#` | `jobs/j1` | `jobs/__dlq/j1` |\n| `iot/+/work/+` | `iot/site/work/j1` | `iot/site/work/j1/__dlq` |\n\n`saguin-dlq-channel` names the queue it came from, so a redriver holding\nthe configuration can look the filter up rather than guessing from the\ntopic.\n\n**Publish the payload back to that topic**, carrying the publisher's own\nproperties - Content Type, Payload Format, and any User Property of its\nown.\n\n**The failure metadata cannot ride along, and that is the broker's job\nrather than the redriver's.** `saguin-` is a reserved prefix on a publish:\nSagüin drops every property in it from any client, so a redriver that\ncopies the whole property set back without thinking still cannot produce a\njob carrying `saguin-dlq-reason`. A publish setting one reaches its\nconsumer without it, whoever sent it.\n\n**`saguin-id` is the one reserved name a publisher may set**, and passing\nthe dead letter's back is the whole reason to look at properties at all.\n\n**What it costs, said rather than discovered.** The redriven job is a\n**new record**: a new offset in the queue, and a fresh attempt count that\nstarts at 1. What survives is its identity - the broker keeps a\npublisher's `saguin-id` (invariant 8), so a consumer can deduplicate across\nthe whole round trip and tell that this is the same work rather than new\nwork that looks like it. Pass nothing back and the job gets a new identity\nand is indistinguishable from work that had never failed.\n\n**And the hazard, out loud.** Redriving into a queue whose bug is not\nfixed dead-letters the job again, with a second set of `saguin-dlq-*`\nvalues and one more offset in both channels. Doing that in a loop is a way\nto spend an afternoon and fill a disk.\n\n### Restart\n\nNo record is in flight after a restart. Deliveries, Delivery IDs, and\ndeadlines are not restored: a restored deadline belongs to a session that\nno longer exists, and a restored Delivery ID could be resolved by a client\nretrying across the very restart that interrupted it.\n\nEvery previously held record is available immediately. Attempt counts\nsurvive, so a record that had already exhausted its attempts is\ndead-lettered rather than starting over.\n\n## Client-declared partitioning\n\n**A subscriber may take a slice of what its filter reaches, chosen by\nitself, with per-topic ordering intact and no coordination between\nmembers.** It declares one MQTT 5 User Property on its `SUBSCRIBE`:\n\n```\nsaguin-filter topic_hash(3, 0)\n```\n\nThe broker delivers a message only when\n`topic_hash(topic) mod partitions == index`, where `topic_hash` is FNV-1a-64\nover the topic followed by the mixing step *The hash, in full* writes out.\n**Declare nothing and you get everything**, which is every client that has\nnever heard of this.\n\n**A member wanting several slices repeats `saguin-filter`, and the calls\nare an OR.** MQTT allows a User Property to appear more than once, which is\nmore idiomatic than a comma-separated list. `topic_hash(8, 1)` beside\n`topic_hash(8, 5)` is a member holding two slices of eight - a member\ncovering a failed peer's share, most often. Asking twice for one slice\nmeans what it says and is not an error.\n\n**Every call on one `SUBSCRIBE` must name the same number of partitions.**\n`topic_hash(4, 0)` beside `topic_hash(6, 1)` is expressible and refused: a\nsubscription has one partition space, which is what the operations listener\nreports and what the broker compares when it warns that two members\ndisagree about the size of a space.\n\n**The arguments are positional, and the order cannot be got wrong\nsilently.** An index must be below the partition count, so if\n`topic_hash(8, 1)` is legal then `topic_hash(1, 8)` is not - which holds\nfor every valid pair, so a client that swaps them is refused rather than\nserved a slice it did not ask for.\n\n**The property is on the packet, not on a filter**, as `saguin-deletions`\nis, so a declaration applies to every filter named; different slices for\ndifferent filters are different `SUBSCRIBE` packets.\n\n**MQTT 3.1.1 has no User Properties**, so a 3.1.1 subscriber cannot declare\na slice and always receives everything its filter reaches. That is a\nprotocol limit rather than a policy, and it is said here rather than left\nto be discovered.\n\n### Where it applies\n\n| | |\n|---|---|\n| broadcast | allowed, on live traffic and on the retained pass a `broker.retained` store hands a new subscriber |\n| `append` | allowed, on the replay from a stored position as well as on live records |\n| `latest` | allowed, on the pass of current state as well as on later changes |\n| `$share/…` | **refused** - a shared subscription already divides a stream between its members |\n| `$saguin/queue/\u003cname>` | **refused** - a queue already divides work between its workers |\n\nBoth refusals are `0x83`, with the reason naming which rule was broken.\nThey are decided by the two prefixes as strings: nothing here asks which\nchannels a filter reaches, because that question is the one *What a filter\nreaches* exists to avoid asking at subscribe time.\n\n**`latest` takes the predicate on both halves or neither.** A member sent\nthe whole of current state on subscribe and then only its share of the\nchanges holds a copy of the key space that starts complete and drifts,\nwhich reads as correct on the day it is set up and as loss a week later.\n\n**The same rule reaches broadcast, and for the same reason.** A configured\n`broker.retained` store hands a new subscriber the current value of every\ntopic its filter reaches, which is a current-state pass however different\nit looks from a channel's - so it takes the predicate too. Broadcast needs\nno channel, no storage and no position, and the predicate still has to\nreach that one handover: a group whose members each received every retained\ntopic is state processed once per member rather than once per group.\n\n**A declaration applies to every filter in its SUBSCRIBE, and to no\nsubscription outside it.** A client holding a declared filter and an\noverlapping undeclared one is served the union: the undeclared\nsubscription is owed everything it matches, and a record reaching this\nclient through either is delivered once. What a declaration narrows is the\nsubscription it was made on, never the client.\n\n### The hash, in full\n\nThe hash is **FNV-1a, 64-bit, over the topic's bytes, followed by a\nmixing step**, and it is a constant of the protocol rather than an\nimplementation detail: a client that wants to know which of its topics land\nin its slice must compute the same value. Both halves are written out here\nso that they can be reimplemented from this document alone.\n\n```\noffset basis = 14695981039346656037\nprime = 1099511628211\n\nhash = offset basis\nfor each byte b of the topic, in order:\n hash = hash XOR b\n hash = hash * prime (modulo 2^64)\n```\n\nThe bytes are the topic's UTF-8 bytes, not its characters, and the\nmultiplication wraps at 64 bits.\n\nThen the mixing step, on that value, with the same wrapping:\n\n```\nhash = hash XOR (hash >> 30)\nhash = hash * 13787848793156543929 (modulo 2^64)\nhash = hash XOR (hash >> 27)\nhash = hash * 10723151780598845931 (modulo 2^64)\nhash = hash XOR (hash >> 31)\n```\n\n`>>` is an unsigned shift right. The slice is taken from the value after\nthis step, never from the value before it.\n\n**Both halves are given in the worked examples below**, so that an\nimplementation that disagrees can tell which of the two is wrong. An\nimplementation that stops after FNV-1a agrees with the *FNV-1a* column and\nwith nothing else.\n\n| Topic | FNV-1a-64 | after mixing | `mod 3` | `mod 8` |\n|---|---|---|---|---|\n| `iot/depot/events/dev-1` | 9321193355118713112 | 16961261177379703000 | 1 | 0 |\n| `iot/water/w-7/inspect` | 14147658934179859562 | 16406880391913018636 | 2 | 4 |\n| `orders/resize/thumbnails/42` | 15020795679369718730 | 10596179872049748585 | 0 | 1 |\n| `a` | 12638187200555641996 | 198367012849983736 | 1 | 0 |\n\n**Why the mixing step is there, and why it is not optional.** FNV-1a stirs\nthe top of its accumulator thoroughly and the bottom hardly at all, and the\nbottom is the half a modulus reads. Flip one bit of a topic and FNV-1a's\nlowest bit changes 12.5% of the time where an even split needs 50%; the\nnext three change 19%, 25% and 31%. That is not noise that averages out.\nBecause the prime is odd, each byte either flips the lowest bit or does\nnot, so a byte appearing twice cancels itself out - and a topic scheme that\ncarries an identifier twice, such as\n`devices/\u003cid>/messages/devicebound/\u003cid>`, sends every identifier to the\nsame low bits. Four members splitting 4000 such topics receive\n2454, 0, 1546 and 0 of them: two members are sent nothing at all, at any\nscale, and *Coverage is the client's responsibility* means nothing reports\nit. With the mixing step the same 4000 split 1005, 996, 1011, 988.\n\n**The step is the ordinary shape of a hash rather than a repair to this\none.** Every modern hash is a mixing loop followed by a finalizer; murmur3\nand xxhash both end this way, and FNV-1a is the unusual one for stopping\nearly. The two constants are splitmix64's, and after the step every output\nbit lands within half a point of the ideal 50%.\n\n**Why FNV-1a and not something else**: xxhash needs a library in every\nclient language, murmur3 has incompatible variants people get wrong, and\nSHA-256 truncated - the honest runner-up - is around twenty times the\nwork per topic for a distribution these ten lines already reach.\n\n**Explicitly not a randomly seeded hash** such as Go's `maphash`: it is\nfaster than all of them and reshuffles every assignment on a restart, so a\nclient could never predict its own slice.\n\n### What a subscriber may declare\n\nA `saguin-filter` value is a call, and `topic_hash` is the only function\nSagüin reads:\n\n```\ntopic_hash(\u003cpartitions>, \u003cindex>)\n```\n\n| | |\n|---|---|\n| `\u003cpartitions>` | 1 to 2147483647, and the same on every call of one `SUBSCRIBE` |\n| `\u003cindex>` | 0 to `partitions - 1` |\n| the property | repeated for several slices, which are an OR |\n\n**An argument is one or more decimal digits**, `0` to `9` and nothing\nelse - no sign, no `0x`, no exponent - so `topic_hash(+3, 0)` is refused\nfor its spelling, and leading zeros mean what they say. **Whitespace is\nthe ASCII four alone**, ignored around the name, parentheses and comma; a\nnon-breaking space is not a space, stated because a value one broker\ntrims and another does not is a value the two answer differently.\n\nNothing else in the value is ignored: the spelling is exact, and\n`TOPIC_HASH` is not this function for the same reason `$SHARE` is not a\nshared subscription.\n\nAnything else is `0x83` with the reason naming the property and stating the\nform: a value that is not a call, a function Sagüin does not have, the\nwrong number of arguments, an argument that is not decimal digits, an\nargument above the bound, a partition count of 0, an index at or above the\ncount, or two calls naming different counts. **An argument that is not a\nnumber and one that is too large get different sentences**, because they\nsend a reader looking in different places: told that 10^23 is \"not a\nnumber\", somebody goes hunting for a stray character that is not there.\n\n**The sentence travels on the SUBACK, as its Reason String**, rather than\nonly in the broker's log: the person who has to act on it is the author of\nthe refused client, who cannot read the log. MQTT's own rule is the one\nexception - a client that connected with Request Problem Information 0 has\nasked for no such text and is sent the reason code alone.\n\nAn empty value is refused rather than ignored - a client that asked for a\nslice and was silently served the whole channel would have no way to\nnotice. A count of 1 is legal and degenerate - one slice, holding\neverything.\n\n**There is no operator syntax.** A client writes a function name and its\narguments, so `$topic_hash % 3 == 0` and every other expression is refused\nrather than read. A name and two arguments admit exactly what the broker\nimplements, and a further predicate is a further function name rather than\na larger language.\n\n**The upper bound is 2147483647 because that is a number every build\nagrees on**: Sagüin runs on 32-bit gateways as well as 64-bit servers,\nand a threshold taken from the platform's integer would make the same\n`SUBSCRIBE` legal on one and refused on the other.\n\n### What this does not build, and what that costs\n\nThere is no group object, no membership, no generation ids, no rebalancing\nand no group cursor. Each member stays an ordinary session keeping its own\nposition, exactly as it does without a declaration.\n\n**Coverage is the client's responsibility, and the broker cannot police\nit.** Nothing stops one member declaring count 3 index 0 while another\ndeclares count 2 index 0. A slice nobody claims is records nobody receives,\nand every member looks healthy. The broker does not report a gap, because\nit cannot tell one from a member that has not connected yet - any gauge\nclaiming to would be wrong during every rolling restart.\n\n**What it does report is two subscribers on one filter declaring different\ncounts.** That is not a timing artefact and cannot be correct, so it is a\nlog line naming both clients. It is scoped to the identical filter string,\nso it misses overlapping-but-different filters - `iot/+/events` against\n`iot/site-a/events` - and that miss is deliberate: where the filters differ\nthe counts may differ on purpose.\n\n**A member's position advances past records it never received.** The\nposition is one cursor into the channel, so a record outside a member's\nslice is stepped over exactly as a record matching none of its filters is,\nand the cursor moves past it. Two consequences follow and neither is a\ndefect:\n\n- **Widening a declaration later does not recover what was skipped.** Those\n records are behind the consumer's position. A member that takes a larger\n slice tomorrow receives more of what arrives from tomorrow, and nothing\n of what it passed over yesterday.\n- **A slice nobody claims is not held anywhere.** Those records age out\n under retention like any other, on the channel's own schedule.\n\nThat is the cost of having no group cursor, and it is the trade this design\nmakes deliberately: the 80% of a consumer-group protocol that is not paid\nfor.\n\n**A declaration lasts as long as the session that made it**, and survives a\nreconnect that resumes one. MQTT restores a resumed session's subscriptions\nbut not the `SUBSCRIBE` that made them, so a member that had to re-declare\nwould be served the whole channel - every other member's slice included -\nbetween resuming and subscribing again.\n\n**That is the opposite of the choice a `latest` channel's deletions make,\nand the asymmetry is which way each fails.** A subscriber that must ask for\ndeletions and has not is served *less* than it wanted: it misses them.\nA member that must re-declare and has not is served *more*: the whole\nchannel, including every slice belonging to somebody else. One degrades\nsafely and the other duplicates records across a group, so they are not\nmade to behave the same way for symmetry's sake.\n\nA session that is *discarded* rather than resumed takes its declarations\nwith it, so a declaration never outlives the session it was made on.\n\n## Ordering\n\n| | Guaranteed |\n|---|---|\n| `append` | total order by offset; each subscriber receives records in that order |\n| `latest` | per topic. No ordering between different topics |\n| `queue` | records are *offered* in offset order |\n| broadcast | per subscriber, for what that subscriber receives |\n\nFor a queue, offered order is not completion order and cannot be. With two\nworkers, or one redelivery, this is normal:\n\n```\nrecord A ─▶ worker 1\nrecord B ─▶ worker 2\nA times out\nrecord A ─▶ worker 2 ← after B is already done\n```\n\nAn application needing strict ordering needs one worker, and even then a\nredelivery reorders against the records that passed it. `append` is the\nprimitive with an ordering guarantee.\n\n## What Sagüin does not guarantee\n\nStated together, because each is a question someone will otherwise have to\nfind out by being surprised:\n\n- **Exactly-once processing.** At-least-once, always. Consumers must be\n idempotent.\n- **Duplicate suppression at publish.** Two publishes with the same\n `saguin-id` become two records in v0.1.\n- **That an acknowledgement was applied.** The worker is not told.\n- **Completion order on a queue.** See above.\n- **That a retained message on a broadcast topic is kept without bound.**\n The retained store keeps it inside its provider's `max_bytes`, and the\n value dies at the publisher's expiry or the operator's\n `retention_period`, whichever runs out first (\"Retained messages\"\n below).\n- **That a broadcast reached anyone.** A broadcast is delivered to whoever\n is subscribed at that moment, and stored nowhere unless it asked to be\n retained. `PUBACK` says the broker accepted the packet - at QoS 1 or 2,\n that every session owed it has kept it (\"Broadcast\") - and nothing about\n who received it, so a publish to a topic nobody holds is `0x00` -\n including a channel name with a typo in it, which is an ordinary broadcast\n topic as far as the broker can tell.\n- **That memory-backed data survives a crash.** Only a graceful shutdown,\n and only as far as the last successful snapshot.\n- **That a Will the broker accepted can still be delivered when it fires.**\n A Will whose topic could never work is refused at CONNECT, where the\n device is there to be told. What cannot be known until it fires - a full\n channel, a storage failure - is logged and no more, because the client is\n gone by then and there is nobody to answer.\n\n## Retained messages\n\n**The retain flag asks for a store, and Sagüin honours it wherever it has\none.** There are two stores, and the topic decides which of them applies or\nwhether neither does:\n\n- **A topic an `append` or `latest` channel claims.** The channel is the\n store. On a `latest` channel the flag is redundant - that channel holds\n the current value of every topic beneath it whether or not anyone set it -\n and on an `append` channel the record is stored as any other and\n replayed to whoever reads the channel afterwards. Nothing is refused and\n nothing extra happens.\n- **A topic a `queue` claims: accepted, and the flag dropped.** A queue\n holds work rather than state, and nothing on one could ever be sent a\n retained value - so the flag is the only part of the publish that\n cannot be honoured, the message becomes a job, and the `PUBACK` says\n the work was stored, which is true. `saguin_queue_retain_ignored_total`\n counts it, nothing on the wire saying a flag went missing.\n- **A broadcast topic.** The retained store holds it (RFC 0002 \"Retained\n messages on a broadcast topic\"). One current value per topic, a\n zero-length payload deleting the entry, handed to whoever subscribes\n next with the RETAIN flag set, and delivered live to whoever is\n subscribed at the time as well, because it is still a broadcast. A value\n sent on subscribe at QoS 1 or 2 to a session that outlives its\n connection is kept in the broadcast log for that session alone, as a\n copy no other subscription or group is owed, so a restart sends it\n again under its identifier, with the RETAIN flag, until it is\n acknowledged (MQTT-4.4.0-1).\n\n**A bridge forwards the live fan-out and never a stored pass.** An outbound\nbridge rule carries what is published while it is running; it is not sent\nthe retained values a subscribe earns, nor a `latest` channel's current\nstate. Two reasons, and either alone would settle it. A stored pass is\nserved on every subscribe, so a bridge that took one would ship the whole\nof current state again at each reconnect - deterministic duplication at the\nfar end, growing with every link drop. And a value a bridge itself brought\nin is still in that store, so forwarding one would hand a peer its own\nretained set back, once per restart, for ever. The record keeps the bridge\nit arrived on for that second reason, in the retained store as everywhere\nelse.\n\n**A link that drops is not a rule that stopped.** An outbound rule runs\nbeside its link rather than inside a connection, so it goes on taking a\n`latest` channel's changes while the link is down, and each topic's newest\ncrosses when the link comes back - a newer value replacing one still\nwaiting, as it does for any subscriber. What a peer is not sent is what was\npublished before the rule started, which is the same answer a broadcast\ntopic gets, **and a value the channel's own retention removed while it\nwaited**: a topic that went quiet for longer than `retention_period`, or\na deletion older than `deletion_retention_period`, is let go by the rule\nas the channel lets it go, and never crosses. A stale reading is worse than\nnone at the far end as it is here (\"What a reconnecting consumer receives\").\n\n**That mark is what stops the echo, and it carries the weight because the\nflag crosses.** A retained value arriving over a link is kept at this end,\nand a `direction: both` rule reads what arrives - so the only thing between\na retained value and a pair of brokers handing it back and forth is the rule\nthat a record which came in over a bridge is never sent out over one. It is\na blanket rather than a comparison against the bridge's own name, because a\nname permits a second hop and a ring of brokers would each see a name that\nis not theirs. What an operator does instead is re-origination: a broker\nmeant to relay publishes the record again as its own.\n\n**The RETAIN flag crosses as its publisher set it, in both directions.** A\nrecord published retained here is carried to the peer retained, so the peer\nkeeps it and a subscriber arriving there afterwards is served it. Inbound is\nthe same: the bridge subscribes at the peer asking for Retain As Published,\nbecause MQTT puts the flag down on a live delivery otherwise\n[MQTT-3.3.1-12], and a value that reached the bridge looking ordinary would\nbe stored here as ordinary.\n\nThat is the crossing publish only. The paragraph above still holds whole: a\nbridge ships what is published while it runs and never a stored pass, so a\nlink restart re-ships no stored set and neither end's retained set is handed\nback to the other. What a link restart *can* repeat is an unacknowledged\nrecord the upstream re-sends at session resume, which MQTT requires of it;\nRFC 0002 \"Bridges\" has what the bridge does about that and where it stops.\n\n**Retained-ness therefore extends one hop**, which is what \"publish\nretained, and whoever subscribes later sees current state\" means across a\nlink, and it is what mosquitto's bridge already does - so a topology moved\nfrom mosquitto behaves as it did. A `latest` channel at the receiving Sagüin\nremains the stronger form of the same promise: current state per topic,\ndurably, bounded, and readable by a point read rather than only on\nsubscribe.\n\n**Where it does not extend, and this is the degradation to know.** MQTT 5\nlets a broker advertise that it does not accept retained messages, and a\npeer that does is sent the value with the flag down: its subscribers receive\nit live and a subscriber arriving there later finds nothing. The alternative\nis worse and is not offered - a retained publish to such a peer is refused,\nand a bridge that carried the flag regardless would hold that record, retry\nit on every reconnect, and never move again. So the value crosses and only\nthe retained promise is lost. Sagüin says so once per link in its log rather\nthan once per record. Where the peer can advertise the limit (MQTT 3.1.1\nhas no way to), Sagüin honours it without being configured to.\n\n**A value the receiving broker will not keep is refused as any publisher's\nwould be.** A bridge has no route of its own into a retained store: a\ncrossing record takes the ordinary publish path and meets the ordinary\nanswers - the store's bounds, the `acl_file`, a zero-length payload as a\ndeletion.\n\n**On a broadcast topic the rules are MQTT's exactly, with nothing added.**\nThat is the whole point of it: an unmodified Home Assistant, Zigbee2MQTT or\npiece of firmware works, and anything Sagüin invented here would be\nsomething to discover the hard way. So, in the six places a broker is\ntempted to differ:\n\n- **The `SUBACK` is written first, then whatever the subscription\n earns.** MQTT permits either order (§3.8.4), so this is a choice rather\n than a rule, and Sagüin makes the one every client library is written\n against: a client that reads one packet expecting its acknowledgement\n gets one. It holds for everything a SUBSCRIBE brings back, a channel's\n replay and a `latest` channel's current state along with the stored\n values here. The subscription itself is recorded before the `SUBACK`,\n so a message published in the gap is delivered rather than missed.\n- **Delivery happens at SUBSCRIBE, not at connect.** A client that resumes\n a session and does not subscribe again is sent nothing, exactly as a\n standard broker does. The difference from a `latest` channel is\n deliberate and is described there.\n- **Retain Handling is honoured**: send the stored values at subscribe,\n send them only if the subscription did not already exist, or send none.\n A client asking for none and receiving them anyway is a broker breaking\n its own word.\n- **Nothing is sent to a shared subscription**, which is what MQTT says of\n retained messages, and Sagüin adds no exception. A shared subscriber is\n served what is published while it is there and nothing that was stored\n before it arrived.\n- **The RETAIN flag arrives as MQTT says it should.** A subscriber served\n the stored value at SUBSCRIBE sees it set. A subscriber that was already\n there sees it cleared on the live delivery, because that is an update\n rather than state - unless it asked for Retain As Published, which is a\n client saying \"tell me what the publisher set\", and then it sees it set.\n Sagüin's own store is still the only one there is: nothing is left in the\n substrate's.\n- **The publisher's Message Expiry Interval discards the stored value**\n (MQTT-3.3.2-5), so the value dies at the earlier of that interval and\n the operator's `retention_period`, each clock deleting on its own. A\n subscriber arriving past the expiry is not served it, even before the\n sweep has reclaimed its bytes, and one arriving inside it is served the\n remaining countdown - the publisher's, decremented by the wait, and\n never the operator's period, which no client hears about. A value whose\n publisher set no expiry carries no countdown and waits on the\n operator's clock alone. Both deletions are silent, as retention is\n everywhere. This is the one MQTT rule that stops at the channel\n boundary: on a channel the record outlives its expiry, and only the\n countdown says so - the split is under \"From publish to record\".\n\nA value goes back out as it was published, carrying no `saguin-id`, no\n`saguin-offset` and no `saguin-timestamp`. So \"how old is this state\" is\nnot answerable here and is answerable on a `latest` channel, which is one\nof the reasons to name a channel for a prefix that matters.\n\n**A client whose roles deny `retained` has a retained publish to a\nbroadcast topic refused with a `DISCONNECT` carrying `0x9A` (Retain not\nsupported)** (RFC 0002 \"Taking a feature away\"). It is a disconnect rather\nthan a reason code because MQTT makes it a protocol error: `0x9A` is not\namong the codes a `PUBACK` may carry, and at QoS 0 there is no `PUBACK` at\nall. The alternative - dropping the flag and acknowledging - tells a\nproducer its state was kept when nothing kept it, and that is the failure\nthis whole document exists to prevent. The denial is about broadcast topics\nonly: a retained publish to a channel's topic is taken as it always is, and\nthat client's `CONNACK` still says Retain Available 1.\n\n**What MQTT's own retained store lacks, both of Sagüin's have.** A broker's\nretained store is conventionally broker-wide, unbounded, without retention\nand without configuration, which is what invariant 13 refuses. Sagüin's\nsits on a named provider, is bounded by that provider's `max_bytes`,\nrefuses rather than evicts at the bound, survives a restart to whatever\ndegree its provider promises, and can be read with `sqlite3` while the\nbroker runs. That is the whole of the difference.\n\n**Which of the two you want.** A `latest` channel is the answer when the\ntopics are a fleet's state and you want them bounded, replayable as current\nstate, and named in the configuration file. The retained store is the\nanswer when a client insists on the flag and you would rather it worked\nthan have to name every prefix it uses - a device tree published under\n`homeassistant/…` or `zigbee2mqtt/…` works either way, and works with no\nSagüin-specific configuration at all.\n\n**The `CONNACK` carries Retain Available 1** to every client, a client\ndenied `retained` included, because MQTT has one answer for the whole\nbroker and no way to say \"on these prefixes\".\n\n## Last Will\n\n**A Will is delivered through the ordinary publish path**, so it may target\nbroadcast or any channel type and gets exactly what a publish gets: the\nreserved `saguin-` prefix stripped from its headers, a Message ID and an\noffset, storage in the channel that claims its topic, and every bound\napplied.\n\n**Including the publish properties it carried.** A Will may set a Content\nType, a Payload Format Indicator, a Message Expiry Interval, a Response\nTopic and Correlation Data, and every one of them is delivered with it -\nalongside its User Properties, and on a channel stored with the record like\nany other publish's. That is the same list \"From publish to record\" gives\nfor an ordinary publish, which is the point: a Will is a publish whose\nsender has gone, and a monitor reading one should not have to know it was a\nWill to know what the bytes are or where to answer. A Will into a `queue`\nis an ordinary job with a Delivery ID and a Response Topic, which is a\nreasonable thing to want when a device dies.\n\nThat is not how the substrate publishes one - `sendLWT` calls\n`publishToSubscribers` directly and never `OnPublish` - so a Will left to it\nreaches a channel's live subscribers while being stored nowhere, hands a\nqueue worker something no timeout will ever reclaim, and skips the bounds\nentirely. Routing it is what makes a Will an ordinary record instead of\none more special case.\n\n### What is refused, and when\n\n**At CONNECT, with `CONNACK 0x90`, for a Will that could never be\ndelivered:** a topic in the reserved `$saguin/` space, a topic in a\ndead-letter channel - which takes records from its queue and from nothing\nelse - a topic longer than `max_topic_length` or deeper than\n`max_topic_levels`, which nothing else bounds, since a Will topic arrives in\nCONNECT and is not checked again until it fires, and **a topic name nothing\nmay publish to at all**: one holding `+` or `#`, one in the `$SYS` tree the\nbroker keeps about itself, or one beginning `$share/`, which is a\nsubscription form (RFC 0002 \"Publishing\").\n\nThat last one is asked here because the answer is worthless anywhere else.\nA wildcard in a topic name is refused a publisher on the wire by the\nprotocol itself, but a Will arrives inside CONNECT where no such check\nruns - and a Will is the one thing a client can leave behind that outlives\nit. Stored, it would sit at the head of the channel for ever: MQTT forbids\na wildcard in a *delivered* topic name as firmly as in a published one, so\nevery conforming consumer stops on reaching it, and on a `latest` channel\nno client could remove it, because deleting a value means publishing an\nempty payload to its topic and that publish is refused too.\n\nRefusing costs the connection, MQTT requiring a server to close after a\nnon-zero CONNACK - **and that is the point**: each of these is permanent\nfor that client's configuration, and the alternative is a device that\nconnects happily for months while its fleet monitoring was never armed.\n\n**At CONNECT, with `CONNACK 0x87`, for a Will on a topic its client's\nroles may not publish to** (RFC 0002 \"What a client may do\"). A Will is a\npublish the client asks the broker to make for it, so it meets the rules\nthe client's own `PUBLISH` meets. Sagüin publishes it as itself, which no\nrule is asked about, so without this a device whose role only reads a\n`latest` channel could overwrite that channel's state by dropping its link.\n\n**A retained Will follows the retained rule and nothing else.** It is\nadmitted wherever a retained publish to the same topic would be - an\n`append` or `latest` channel, a queue where the flag is dropped, or a\nbroadcast topic - and refused at CONNECT with `0x9A` where a publish is\nrefused: a broadcast topic from a client whose roles deny `retained`. That\nit is the same rule is the point of it, since a device that may publish a\nmessage all day should not be turned away at CONNECT for promising to send\nthe same one. A retained Will aimed at a queue is still the availability\npattern aimed at the wrong channel type - a device means \"tell a dashboard\nI died\", and one worker takes it once and is done - but that is the\noperator's `filter` saying the topic is work, and refusing a connection is\nnot how they find out. A retained Will into a `latest` channel is the\nstandard availability pattern, a retained `online` and a retained Will\npublishing `offline`, and it works.\n\n**One question is asked again when the Will fires: may its client still\npublish there.** The `acl_file` can be reloaded between a `CONNECT` and a\nWill that fires days later, so the Will is judged as the client that armed\nit - its client id and the name it authenticated as, which its session\nkeeps beside it - against the rules as they stand. A Will refused then is\nnot published: it is logged and counted in `saguin_publish_refused_total`\nas `not authorized`, and not in `saguin_wills_published_total`. Nothing\nelse is refused when the Will fires. A full channel or a storage failure\nis logged and no more: the client is gone by then, so there is nobody to\nanswer, which is exactly why everything that can be judged at CONNECT is\njudged there.\n\n### The delay is Sagüin's, not the substrate's\n\nA client may ask for its Will to wait - a Will Delay Interval - so that a\nlink dropping for four seconds does not announce a device that is still\nthere. **Sagüin holds that Will itself**, delivers it when the delay\nexpires, and cancels it when the same client id resumes the session - a\nconnection as \"Sessions\" defines one, never a CONNECT that was refused. A\nClean Start ends the session instead, which publishes it. A Will with no\ndelay is published as its connection ends, a takeover included.\n\n**The wait ends when the session does, whichever comes first**, which is\nMQTT's own rule (5.0 section 3.1.3.2.2) and is two cases rather than one. A\nsession that ends with its connection - Clean Start with no Session Expiry\nInterval, or a 3.1.1 client with Clean Session 1 - ends at the disconnect, so\nits Will is due then however long its interval says: holding it would be\nwaiting to announce the death of a session the broker has already thrown\naway, and a device asking for five minutes would be announced five minutes\nafter nothing was left to announce. A session shorter than the interval ends\nfirst too, and what counts is the interval the session was **granted**: a\nclient asking for an hour on a broker whose `limits.max_session_expiry` is a\nminute has a minute of session, and its Will is due at the end of it.\n\nIt is held by Sagüin, and the substrate holds no Will for later, so a\ndelayed Will is published through the ordinary publish path when it falls\ndue, under every rule above - and a delay is precisely what a flapping\nlink makes an operator configure.\n\n**A delay needs a session to wait inside, and that is the recipe.** A device\nthat wants flap protection asks for a Session Expiry Interval at least as long\nas its Will Delay Interval; one that asks for a delay on a session ending with\nits connection is asking for a wait its own session will not allow, and its\nWill is published at the disconnect. Sagüin says so in its log when it\nhappens, naming the client, the interval it asked for and the fix, because\nnothing on the wire can carry that answer and the device is gone by the time\nit would matter.\n\n**The wait never outlives the session.** Sagüin publishes the Will at the\ndisconnect where the session ends with it, and caps the wait at a real\nsession expiry. A Will outliving its session is the one thing Sagüin's\nstorage rule forbids, because a Will belongs to a session, is discarded when\nit ends, and therefore needs no home, bound or restart rule of its own. A\nfleet migrating from a broker that waits out the whole interval, with Clean\nStart and a delay, will see its announcements arrive at the disconnect\nrather than after the interval; the log line above is where that shows.\n\n**What the wait costs while it runs** is one small record per waiting client:\na client id, the moment the Will is due, and a timer. The Will itself is on\nthe session record, in the provider `broker.session.storage` names, where the\nCONNECT that armed it wrote it - so a waiting Will is held once rather than\ntwice, and `saguin_wills_waiting` is how many are held (RFC 0005).\n\n**A restart does not lose it.** The moment a Will becomes due is written\ndown when the wait begins, so a broker that comes back inside the wait\nresumes what is left of it, and one that comes back after the moment has\npassed publishes then - late by however long the outage was, which is the\nhonest cost of the outage rather than a lost announcement. A start ends\nsessions as well as restoring them - one whose expiry passed while the broker\nwas stopped, one that ended with a connection the stop cut, one holding a\nsubscription the current rules refuse - and where the Will it holds carries\na moment, ending the session is that Will falling due: it is published\nbefore its record goes. **So a Will is published at least once across a\ncrash, and can be published twice**: a Will that falls due is published\nbefore the write that takes it off its record - when its delay runs out,\nwhen its session expires, or at a start - so a broker that dies between\nthe two leaves the Will on its record, where the next start can find it\nowed again and publish it again. The other order would lose it instead. A\nWill with no moment on it is the next paragraph's case, and the two outage\ncases below are what it costs.\n\n**A client that was connected when the broker stopped keeps its Will**, and\nthat Will has no moment on it, because nothing has happened yet to make one: a\nbroker stopping is not a device dying. It waits on the restored session for\nwhichever comes first. The client returns, and its `CONNECT` replaces the\nWill. Or the session expires at the running broker - which is that broker\nwatching the client stay away for the whole interval it was granted - and the\nWill is published then. Stripping it at the start was the other option and it\nis the worse one: every restart would quietly take Will protection away from\nevery connected device until each happened to reconnect.\n\n**A clean `DISCONNECT` withdraws the Will from the session's record**, with\nthe expiry it carried, before the connection that answers it is closed\n(invariant 18), so a crash once the client has seen its connection close\npublishes no Will it withdrew. Where the store refuses that write the broker\nholds the withdrawal itself and writes it again every second until it lands,\nand as it stops - so only a crash while the store is still refusing leaves the\nrecord reading as a connected client's, whose Will is then published when the\nsession expires. A `CONNECT` for that client id writes it first, and is\nrefused `0x83` while the store refuses it: the session it would claim would\nread as though its client had never said goodbye.\n\n**What an outage costs, stated plainly, because an operator will meet it.**\nTwo cases are announced by nobody, and both have the same shape - the broker\nwas not running to see the connection end.\n\n- **A device with a session that ends with its connection**, Clean Start with\n no expiry. Its session is ended at the next start, and its Will goes with\n it: there is no session left for the announcement to belong to.\n- **A device whose session expiry passed while the broker was stopped.** That\n session is ended at the next start too, and its Will carries no moment, so\n nothing says it was due rather than a client that simply stayed away.\n\nBetween them sits the case that *is* covered: a device that goes away, whose\ngranted session outlasts the outage, is announced when that session expires -\nlate by the outage, which is the honest cost rather than a lost announcement.\nA fleet that needs deaths noticed through an outage of any length needs\nsomething outside MQTT watching it, because a Will is a message a broker sends\nwhen it sees a connection end, and a broker that is not running sees nothing.\n"}