From the source of truth

RFC 0002 - Channels and configuration

A static snapshot from the Sagüin repository.

covers what MQTT reserves and nothing else, and every rule that\nfollows from a name then has to be re-derived for whatever is left. A `/`\nis the case that shows why: it would make a name hierarchical, which makes\nit a prefix another name could overlap, a file name that needs encoding, a\nname inside this 128-byte limit still too long to store, and\n`$saguin/consumer/site/events/seek` unreadable as a seek. An allow-list\nanswers all of that once.\n\n`.` and `..` are inside the character set and refused anyway. Both name a\ndirectory rather than a channel in every listing that will ever hold one,\nincluding the snapshot directory where a channel's name becomes a file.\n\nThe 128-byte cap is a constant in the broker, not a configuration key.\nA channel name is written by hand by the operator, once, in a file the\nbroker already trusts - it is not client input, so the limit protects\nnothing and tunes nothing. Exposing it would add a knob whose only\nreachable settings are \"the default\" and \"wrong\".\n\n`max_topic_length` is exposed precisely because the opposite is true of\nit: a topic arrives from an untrusted client on every publish, and an\noperator on constrained hardware has a real reason to lower it.\n\n`max_client_id_length` defaults to 256. A durable consumer's stored\nposition is keyed by its client id, so an unbounded id is durable state\nwhose size the client chooses; a longer one is refused at CONNECT with\n`0x85 (Client identifier not valid)`, before a session exists. MQTT\nrequires a server to accept any id up to 23 bytes, so a configured value\nbelow 23 is refused at startup.\n\n`max_session_expiry` defaults to **30d** and stops a client choosing how\n*long* a stored position lives, as the bound above stops it choosing how\nlarge one is: MQTT's own ceiling is about 136 years, so a fleet taking a\nfresh client id every boot would otherwise leave one row per boot, per\nchannel, that nothing an operator sets removes.\n\nMQTT supplies the whole mechanism, so this is a cap and never a refusal:\nthe server shortens what the client asked for and the CONNACK carries the\nshortened value, so a conforming client is told what it got. A client\nasking for less than the cap is given what it asked for.\n\n**A DISCONNECT is held to it too.** MQTT 5 lets a client change its\nSession Expiry Interval as it leaves (section 3.14.2.2.2), except from 0 to\nanything else, which is a protocol error. The value it sends is the one the\nsession is kept by from then on - across a restart as well - and one above\nthe cap is kept at the cap. A DISCONNECT has no answer to say so in, so the\nclient is not told, as it would not be told of a session ended early for\nany other reason of the server's.\n\nThirty days because too short is the expensive direction: these\ndeployments go offline for weeks at a time, and a consumer whose session\nexpired resumes from the retention floor instead of where it stopped. A\nlong cap never costs correctness (invariant 1), but it does cost memory:\nthe broker holds a session whose client is away for the whole of its\nexpiry, a few kilobytes and more for each filter (\"Every session's state:\n`broker.session`\" has the figures), and nothing else ends one that is\nnever resumed. A memory provider's `max_bytes` bounds how many are held.\nA sqlite provider's `max_bytes` bounds its file instead, so on sqlite - and\non a memory provider with no `max_bytes` - how many are held is bounded by\nthis cap and the rate at which new client ids arrive: a fleet taking a fresh\nid every boot holds every boot's session for thirty days.\n\n`none` is accepted by the retention keys and refused here, because a\nsession that never expires is the unbounded table the key exists to close.\n\n### Which channel a topic belongs to\n\nFilters overlap, and that is the point rather than a mistake to refuse. A\ndomain holds a location, a reading and a command at once, and saying so\ntakes two filters that both match some of the same topics:\n\n```\nwater-location iot/water/+/location/#\nwater-measurement iot/water/+/+/#\n\niot/water/w-7/location matches both\n```\n\n**The filter that spells the topic out most exactly holds it.** Compare\nthe two filters level by level from the left. The first level at which\nthey differ decides, and at that level:\n\n| | |\n|---|---|\n| a spelled-out level | beats `+`, and beats `#` |\n| `+` | beats `#` |\n| a filter that ends here | beats one continuing with `#` |\n\nAt level 3 above, `location` beats `+`, so `iot/water/w-7/location` is a\n`latest` value; `iot/water/w-7/flow` matches only the second filter and is\nan append record. Nothing further needs stating, because the rule reads\noff the two lines the operator wrote.\n\n**The order of the file has no effect on any of it.** Sorting `channels:`\nalphabetically, or moving half of them into a file named by `!include`,\ncannot change where one topic lands. That is deliberate, and it is why\nplacement is a key on the channel rather than a list somewhere else: a\nlist has an order, an order in a file is a thing the next reader changes\nwithout meaning to, and routing that moves when somebody tidies a file is\nexactly the silent class of defect this broker is built to refuse.\n\n**Two channels may not carry the same filter**, and that is a startup\nerror naming both - the only tie the rule can be handed, since two\nfilters equally exact at every level of one topic are character for\ncharacter identical. Resolving by file order gives an order nobody can\nsee once `channels:` is split across files, and refusing overlap outright\nrefuses the three-line example at the top of this document.\n\n### Changing a filter\n\n**Records stay where they landed.** Editing a filter changes where new\nrecords go and moves nothing already stored: a record written under\nyesterday's filter is served from the channel it is in. Somebody will\nchange a filter and go looking for yesterday's records in the new place,\nwhich is why this is written down rather than left to be worked out.\n\nThat is safe in one direction and worth checking in the other. A topic\nthat *starts* reaching a channel is straightforward - new records go\nthere, and a broadcast retained value whose topic now reaches a `latest`\nchannel is moved in at startup (\"Retained messages on a broadcast topic\").\nA topic that *stops* reaching one leaves what was stored behind it\nreachable by nothing: the records are all there, `$saguin/kv/get` answers\nthat the topic names no latest channel, and no counter anywhere is wrong.\n\n**So at startup the broker counts the values in each `latest` channel\nwhose topic that channel's filter no longer matches, and logs the\nnumber.** One value per topic is what makes that cheap. **Append channels\nare not checked**, because it would mean reading the whole log - and that\nis said here rather than left for an operator to conclude from silence\nthat there is nothing to find. `saguin --route` on the old topic answers\nthe question for either kind, before or after the change.\n\n## What a filter reaches\n\n**A subscriber is served every channel its filter touches, and each record\nunder the semantics its own topic has.** A filter touches a channel when\nthere is any topic both it and the channel's filter match. The client\nnever has to know where the channel boundaries are, which is the whole\nreason placement is a filter and not a name.\n\nWith the three channels from the top of this document:\n\n| Filter | Is served |\n|---|---|\n| `iot/water/w-7/location` | that device's current location, carrying RETAIN |\n| `iot/water/+/location/#` | every device's current location |\n| `iot/water/w-7/#` | the device's location as current state, its readings replayed from the subscriber's own position, and anything else beneath it live |\n| `iot/weather/st-3/data/raw` | live broadcast only - no filter claims that topic |\n| `#` | all of the above, across every device |\n\n**A filter spanning two durable channels is two deliveries and not one\nmerged stream.** A durable consumer holds one position per channel\n(RFC 0003 \"Sessions\"), so `iot/water/w-7/#` resumes independently in the\nappend channel and is served the latest channel's current values on every\nsubscribe.\n\n**A queue is the exception, and it is the only one.** A filter crossing a\nqueue is granted and served everything else it matches, with the queue's\nrecords left out - so `mosquitto_sub -V 5 -t '#'` at a new broker works,\nwhich is the first thing anyone types, and takes nobody's work while doing\nit. A queue's records are reachable through one subscription form and no\nother, which is the next section.\n\nEvery row above is **accepted**, and so is a filter that matches nothing\nat all: it matches nothing *yet*, and the topic it was written for may be\npublished a second later. The rule is enforced by the broker at\n`SUBSCRIBE`, against the filter the client sent - an SDK that constructs\nfilters correctly changes nothing about it.\n\n### What a wide filter costs, said rather than discovered\n\n`#` with Clean Start = 1 is served every append channel from its\nretention floor, and a client reconnecting twenty times reads it twenty\ntimes. That is what an append channel promises and it is not a defect, but\nit is a surprise if the first thing typed at a loaded broker is\n`mosquitto_sub -t '#'`. `saguin --route \u003cconfig> '#'` answers what a\nfilter would be served before anybody subscribes with it.\n\nThe resource shape is the same sentence from the other side: a filter is\none delivery goroutine per channel it touches, so `#` on a broker with N\nappend channels is N replays into one socket, each bounded by\n`limits.write_timeout` and the client's Receive Maximum.\n\n**A filter is not served only where it fits inside one channel's filter,\nbecause of what that would hide.** It would make `iot/water/w-7/#` a\nsubscription to the device's commands and not its readings, answered\n`SUBACK 0x00` with no indication that half of what was asked for is\nmissing. A client that asked for everything about one device, was told\nyes, and was given half. Refusing loudly and serving fully are both\nhonest; a silent half is not.\n\n## What each channel type admits: the SUBACK codes\n\n### `append`\n\n| | |\n|---|---|\n| Publish | `PUBLISH` to any topic the channel's filter matches, at any QoS. A QoS 2 publish is held and written at its `PUBREL` (RFC 0003 \"Exactly once\") |\n| Subscribe | Any filter touching the channel's |\n| Position | Per `client_id`, per channel |\n| From 3.1.1 | Publish and subscribe both, with a durable position on `cleanSession = 0`. Where a reader with no position starts is `start` below, and is the same answer for every protocol |\n\nA subscriber with Clean Start = 0 and a non-zero Session Expiry Interval\nis a **durable consumer**: its position is stored and it resumes there on\nreconnect.\n\n**The rest is delivery semantics and RFC 0003 owns it**: the replay from\nthe retention floor, what `start: tail` changes, the position as a cursor\nper `client_id` per channel, what Clean Start = 1 discards, and why a\nconsumer that changes its filter seeks rather than re-reads (RFC 0003\n\"Where a subscription starts\" and \"Durable consumers\").\n\nA shared subscription over this channel is served new records only, split\nacross the group - *Shared subscriptions* below.\n\n### `latest`\n\n| | |\n|---|---|\n| Publish | `PUBLISH` to any topic the channel's filter matches, at any QoS. A QoS 2 publish is held and written at its `PUBREL` (RFC 0003 \"Exactly once\") |\n| Subscribe | Any filter touching the channel's |\n| On subscribe | The current value of every matching topic is delivered |\n| From 3.1.1 | Publish and subscribe both, and current state arrives with the RETAIN flag exactly as it does for MQTT 5 |\n\nA publication replaces the topic's current value. A zero-length payload\ndeletes it, matching MQTT's own retained-message convention.\n\nCurrent state is delivered carrying the RETAIN flag, so a stock client\ncan tell state it is catching up on from an update that has just\nhappened. RFC 0003 is authoritative on how it is delivered.\n\nWhat a `latest` channel adds over a retained message - including over\nSagüin's own retained store, which is bounded and durable too - is that it\nis named in the configuration file, that every value carries an age and a\nposition, and that a consumer reconnecting is sent only what changed. RFC\n0003 is where those are described. Expiry is configured below as\n`retention_period`, and it deletes the *current value* of a topic that has\ngone quiet - which is the point rather than a defect. There is no size\nbound, for the reason under \"Bounds on what a channel holds\": what grows\nhere is the number of topics rather than a history.\n\n**`deletion_retention_period` is a second clock, for deletions.** A latest\nchannel keeps a deletion - a stored value with no payload, which nothing\nreading the channel is shown - so that a broker mirroring it over a bridge\ncan be told a topic is gone; RFC 0003 has why, and this is how long it\nis kept. A day when the file does not say, `none` to keep it as long as\nthe channel exists. A value and a deletion are not the same thing to\nkeep: a value lives as long as it is the truth, a deletion only until\neverything reading this channel has seen it, so a channel keeping values\nfor ever does not have to keep a row for every device ever decommissioned.\nIt applies to no other channel type and is refused on one.\n\nA shared subscription over this channel is served changes only, split\nacross the group - *Shared subscriptions* below.\n\n### `queue`\n\n| | |\n|---|---|\n| Publish | `PUBLISH` to any topic the channel's filter matches, at any QoS, from an MQTT 5 or a 3.1.1 client alike. A QoS 2 publish is held and written at its `PUBREL` (RFC 0003 \"Exactly once\") |\n| Subscribe | `$saguin/queue/` + the channel's **name** - this form and no other |\n| Subscription QoS | 1 only. QoS 0 on that form is `SUBACK 0x83`; QoS 2 is granted at 1 - MQTT grants the lower of what was asked and what is offered, and a queue's offers are sent at QoS 1 (RFC 0003) |\n| From 3.1.1 | Publish yes, consume no. A delivery carries its Response Topic and Correlation Data as MQTT 5 properties, and 3.1.1 has neither |\n\n**A worker names the channel, and needs nothing out of the configuration\nfile.** A queue called `inspections`, whatever its filter says, is consumed\nthrough `$saguin/queue/inspections` - the same name its seek topic, its\nresponse topic and its dead-letter channel are made of. The form matches\nnone of the queue's topics and is not resolved as a filter: the last level\nis a channel name, compared as a name.\n\n**The broker chooses which worker receives a job**, from its own index of\nwho is subscribed to which queue, and does not delegate that to a matched\nfilter. It already owns leases, offers, redelivery and attempt counts, so\nthe choice sits with the same component as the state it depends on - and\none exact string is one population of workers whatever MQTT would have made\nof it. `$share` is an ordinary shared subscription here and reaches no\nqueue (*Shared subscriptions*).\n\n**`$saguin/` is otherwise a space nothing is delivered under** - a queue's\nform and a seek's reply are the two subscribable topics there, and the\nwhole space is set out under *Publishing*.\n\nA queue's filter carries no `{a,b}`, and that is refused at startup.\nBraces expand into two filters and two ways to place a topic, and the\nambiguity is not something a channel's name should have to answer for.\n\nExactly one subscription form is valid. Every other form is refused. With a\nqueue named `inspections`, filtered `iot/water/+/inspect`:\n\n| Attempted | |\n|---|---|\n| `iot/water/+/inspect` | the channel's filter, which is not the form, and which lies inside the queue |\n| `$saguin/queue/inspections/#` | straddles the work and the response topic |\n| `$saguin/queue/inspections/w-7` | a channel name is one topic level |\n| `$saguin/queue/#` | reaches every queue at once |\n| `$saguin/queue/readings` | names a channel that is not a queue, or none |\n| `$saguin/queue/inspections` at QoS 0 | see below |\n\n**Everything under `$saguin/queue/` that is not a queue's form is refused\n`0x8F`.** MQTT would grant such a subscription happily and Sagüin can never\nfeed it, so the subscriber would be connected and empty for ever. Four ways\nto arrive there, each one keystroke from a correct line: a misspelt channel\nname; a dead-letter channel's name, which is a *name* and not where its\nrecords are; an `append` or `latest` channel's name, same reason; and a\nwildcard or an extra level, which a channel name cannot contain.\n\n**The dead-letter case is why this is a rule and not a warning**: failed\nwork lands in `\u003cqueue>__dlq` and a queue is consumed through\n`$saguin/queue/\u003cname>`, so putting the two together is the obvious\nreading - and a dead-letter channel is an ordinary `append` channel, read\nwith an ordinary filter like any other.\n\n**A plain filter lying entirely inside a queue's filter is refused `0x8F`\ntoo.** A queue admits its form and nothing else, so a subscriber asking for\npart of a queue's topics through an ordinary filter is refused rather than\nserved records that belong to a worker (invariant 11). One that merely\n*crosses* a queue - `#`, `iot/#` - is contained by nothing, is an ordinary\nsubscriber, and is granted: it is served everything it matches except the\nqueue's records. Refusing the crossing form would break\n`mosquitto_sub -t '#'`, which is the first thing anyone types at a new\nbroker.\n\nA dead-letter channel's topics lie inside its queue's filter when that\nfilter ends in `#`, and a reader of them is not a worker asking wrongly.\nThe dead-letter filter is the more exact of the two, so it holds those\ntopics and the reader is granted - see \"The dead-letter channel\".\n\n### Shared subscriptions\n\n**`$share` is MQTT's and Sagüin does not take it.** A shared subscription\nis granted on every channel type and over broadcast alike, and is served\nwhat is published after its group began, split across the group:\n\n| | |\n|---|---|\n| broadcast | everything, split across the group |\n| `append` | new records only, split across the group. No replay, no offsets, and it never moves a stored position |\n| `latest` | changes only, split across the group. No pass of current state on subscribe |\n| `queue` | nothing. Queue records only ever leave through the queue's own form |\n\n**A group is served alike over all three.** One with a member whose\nsession outlives its connection holds what it is owed while its members\nare away, and a member's session ending gives back or counts what the\ngroup had handed it (RFC 0003 \"Broadcast\"; \"What a shared group is owed:\n`broker.share`\" below) - a channel's record exactly as a broadcast.\n\nOver an `append` channel it is not load-balanced *consumption* of the\nlog's history, which would need a single position advanced by concurrent,\nunordered acknowledgements - a consumer-group protocol Sagüin does not\nhave. A queue is the primitive for distributing work. The group holds no\nposition on the channel and never moves one - including the position of\nthe same client's ordinary subscription, if it holds one. And no pass of\ncurrent state on subscribe is MQTT's own rule: a server does not deliver\nretained messages to a shared subscription, and a `latest` channel's\ncurrent state is delivered carrying RETAIN - a subscriber that wants\ncurrent state subscribes ordinarily.\n\n**Each message goes to one member that can take it now**: a session that is\nconnected, whose outbound queue is not full, and - for a delivery at QoS 1 or\n2 - with room in the in-flight window its Receive Maximum gives it. The\nmembers that qualify take turns in the order of their client ids, one turn\nper message, with a rotation kept for each group. A member whose client is\naway is not given the group's share to hold, and neither is a connected one\nthat has stopped acknowledging, and neither is one whose Maximum Packet\nSize the record is over, which is passed over rather than disconnected (RFC\n0003 \"When a record is too large for a subscriber\").\n\n**A message no member can take is never handed to a member that cannot\ntake it.** A group with a member whose session outlives its connection\nholds it; any other drops it, counted as\n`saguin_session_deliveries_dropped_total{cause=\"no_shared_member\"}` (RFC\n0005). An expired session is not a member at all: it leaves every group\nit was in when it expires.\n\n**It is never refused for the channels it reaches.** No ShareName is\nreserved, no `0x9E` is sent for what it reaches, and the broker does not\nwork out which channels a shared filter touches. A shared group whose\nfilter lies inside a queue's is granted and receives nothing from that\nqueue, which is honest rather than lax: a queue's records leave through\nits own form and no other, so there is nothing there for such a group to\nbe given.\n\nFour refusals apply to a shared subscription, and none asks anything\nabout the configuration:\n\n- **It is `$share/\u003cShareName>/\u003cfilter>` and nothing looser** (MQTT 5\n section 4.8.2): a ShareName of at least one character holding no `+` or\n `#`, then `/`, then a topic filter. A filter that begins `$share` again,\n in any case, is refused too: it matches only topics under `$share/`, and\n nothing can publish there. `0x8F`.\n- **The filter inside is refused where it would be refused alone.**\n `$share/g/$SYS/#` is refused `0x87`, as `$SYS/#` is, and so is any filter\n under `$saguin/`, a queue's form and a seek reply included: a queue's work\n goes to the one worker Sagüin chose and a seek reply to one client's\n socket, so neither is ever delivered to a group.\n- **`$share` is spelled in lower case and no other.** `$SHARE/grp/events/#`\n is an ordinary filter under a level called `$SHARE`, and Sagüin refuses\n the spelling rather than interpreting it, `0x8F`.\n- **A 3.1.1 client cannot have one.** Its specification gives `$share` no\n meaning, so it has asked for topics under `$share/`, which nothing can\n publish to. Refused `0x80`, the one failure code a 3.1.1 `SUBACK`\n carries.\n\nEach catches a filter that would otherwise be granted and then match\nnothing for ever.\n\nA client whose roles deny `share` is not refused a filter at all: its\n`CONNACK` says Shared Subscription Available 0, and a `SUBSCRIBE` asking\nfor one anyway closes the connection, `0x9E` (*Taking a feature away*).\n\n**Each is refused with the code that describes it**, and this is every\nreason code a `SUBACK` from Sagüin carries - one per filter, so a SUBSCRIBE\nnaming four gets four answers:\n\n| Reason code | When | To a 3.1.1 client |\n|---|---|---|\n| 0x00 / 0x01 / 0x02 Granted | Granted, at the QoS shown - never higher than the one asked for, and never above the broker's ceiling of 2 (RFC 0001), or of 1 for a client whose roles deny `qos2`. A queue's form is granted at most 1, its offers being sent at QoS 1 | the same |\n| 0x83 Implementation specific error | QoS 0 on a queue's canonical form: there is no transport acknowledgement to start the visibility timeout from. Also, on every filter of the packet: a `saguin-` User Property Sagüin does not read, or a partition declaration that is malformed or self-inconsistent, neither of which names a filter. A valid declaration on `$share/…` or a queue's form (RFC 0003) refuses that filter alone, and the packet's other filters are granted with the declaration applied. And on every filter the packet would add or change, the provider `broker.session.storage` names failing to keep the session's record, or the cursor of a shared group kept with it, other than for room - the session is kept as it was | `0x80`. The store's failure is reachable; the rest is not: a 3.1.1 client is refused the form outright, and has no User Properties to carry the others |\n| 0x87 Not authorized | A filter the client's roles do not allow (*What a client may do*), or one in a reserved space nothing is delivered under - `$SYS/`, or `$saguin/` other than the two topics Sagüin defines a subscription for. Either space inside `$share/\u003cShareName>/`, those two topics included | `0x80` |\n| 0x8F Topic Filter invalid | A string that is not a topic filter at all, `+` or `#` taking less than a whole level (MQTT-4.7.1-1, -2); one with more levels than `limits.max_topic_levels`; a spelling under `$saguin/queue/` that is not a queue's canonical form; a plain filter lying entirely inside a queue's filter; a queue's form asked with No Local, which would withhold a job from the one worker whose client id published it (RFC 0003 \"No Local\"); a queue's form from a 3.1.1 client; `$share` written in any case but the one MQTT defines; or a shared subscription that is not `$share/\u003cShareName>/\u003cfilter>`, or whose filter begins `$share` again | `0x80`, and any filter beginning `$share/` at all - except the string that is not a filter, which closes the connection, 3.1.1 having no code for a malformed packet |\n| 0x97 Quota exceeded | The provider `broker.session.storage` names is at its `max_bytes`, on every filter the packet would add or change - the session is kept as it was. The cursor a shared group is given by its first member whose session outlives its connection is kept in the same write as that member's record, so no room for the cursor is no room for the packet. Also a filter that would take the client past `limits.max_subscriptions` (\"How many topic filters one client may hold\") | `0x80` |\n\n**A refusal is total and the client stays connected**, whichever code it\nis: it is told which rule it broke and may subscribe again correctly - a\n3.1.1 client included, the `SUBACK` being the one packet where 3.1.1 can\nbe refused in words rather than by a closed connection. The exceptions\nare protocol errors rather than filters, and they end the connection\ninstead: a `SUBSCRIBE` carrying no filters at all - `0x82`, under \"A\n`DISCONNECT` from the broker\" - a shared subscription asked with No Local\n(MQTT-3.8.3-4; RFC 0003 \"No Local\"), and a malformed filter from a 3.1.1\nclient, which that protocol gives no code for.\n\n**Except where the substrate granted it anyway**, and then the client is\ndisconnected - `0x9B`, `0x8F` or `0x83` matching the refusal that was\noverridden (RFC 0003) - rather than left holding a subscription Sagüin\nrefused. Every *granted* filter is checked a second time because a hook\nonly advises: a substrate that ignored the reason codes would grant\nexactly what was refused (invariant 10). On one that honours them this\nnever fires, and either way the broker never silently downgrades a queue\nsubscription into something that looks like it worked.\n\n**A queue cannot be consumed by a 3.1.1 client at all** - a delivery\ncarries its Response Topic and Correlation Data as MQTT 5 properties, and\nacknowledgement is a publish echoing both, so such a worker would be\nhanded jobs it can never resolve: each times out, redelivers and\ndead-letters, with nothing reporting that the worker was incapable rather\nthan slow. It is refused `0x80`, the one failure code a 3.1.1 `SUBACK`\ncarries.\n\nA 3.1.1 client publishing *into* a queue is untouched - it is an ordinary\npublish to an ordinary topic, and the whole point of the feature.\n\n**One form is one population of workers** (invariant 4): a second spelling\nis not a second group, it is a subscription that names no queue, and it is\nrefused. Two independent work streams over the same records are two\nqueues; there is deliberately no way to express that with subscriptions.\n\n## Publishing\n\nA publish is resolved to the channel whose filter spells its topic out\nmost exactly, or to broadcast where no filter matches at all. The broker\nreplies with a `PUBACK` (QoS 1) carrying one of:\n\n| Reason code | When | To a 3.1.1 client |\n|---|---|---|\n| 0x00 Success | Stored, or accepted for delivery | `PUBACK`, carrying no code - 3.1.1's has no room for one |\n| 0x80 Unspecified error | On the `PUBREC` of an exactly-once publish from a connection whose client id another connection now holds: the session it would be held for is no longer this connection's | the connection is closed |\n| 0x83 Implementation specific error | The channel's storage, or the broadcast log's, could not keep the record. Also a refused seek - the payload is not an offset or a time, or it is outside the channel - because a seek is a publish to `$saguin/consumer/\u003cchannel>/seek` and this is the answer every client gets, whether or not it set a Response Topic (RFC 0003) | the connection is closed |\n| 0x87 Not authorized | The channel takes its records from one place and this is not it: a derived `\u003cqueue>__dlq`. Also a principal whose roles do not allow it here (*What a client may do*) | the connection is closed |\n| 0x90 Topic Name invalid | Over `max_topic_length`, or deeper than `max_topic_levels`; in the reserved `$saguin/` space without being one of the topics Sagüin defines there; or beginning `$share/`, which is a subscription form rather than a topic | the connection is closed |\n| 0x97 Quota exceeded | The channel is at its size limit and its policy is to reject; the channel's provider, or the retained store's, is at its `max_bytes`; the publish exceeds `max_header_count` or `max_header_bytes`; or the client is over its `publish_rate` or `publish_bytes` (\"How fast one client may publish\") | the connection is closed |\n| 0x99 Payload format invalid | Payload Format Indicator says UTF-8 and the payload is not | cannot arise - 3.1.1 has no Payload Format Indicator |\n\n**A refusal a 3.1.1 client cannot be told about closes its connection, and\nis never acknowledged.** 3.1.1's `PUBACK` carries no reason code, so the\nonly answers the packet can express are success and nothing. Answering\nsuccess to a record the broker threw away is the one failure this project\nwill not ship - a publisher told its record is safe, on a channel\nconfigured to keep it, and the loss visible in nothing but the broker's own\nlog. A closed connection says far less than a code, but what it says is\ntrue, and a client that reconnects and re-sends loses no record. What it\ncosts is a misconfigured device in a reconnect loop, which is at least loud.\n\n**A refused QoS 0 publish is still a silent, counted drop**, on 3.1.1\nexactly as on MQTT 5: there is no reply packet to withhold and nothing\nchanges.\n\nEvery non-success carries a Reason String - but only to a client that\nasked for one: MQTT lets a server withhold it when the client sent\nRequest Problem Information 0, and Eclipse Paho sends 0 by default. That\nis MQTT working as specified, written here because measuring it the wrong\nway round reports a broker that says nothing when it said everything.\n\n**Every refusal is logged whatever the QoS**, so the one case with no\nreply at all is still visible to the operator - and a queue producer has\none more reason to publish at QoS 1.\n\n**A publish at QoS 2 is answered on its `PUBREC`.** Every refusal in the\ntable above is carried on the `PUBREC` instead, and for a reason MQTT\nstates: a receiver makes every check that could produce a forwarding\nfailure *before* it accepts ownership of the message, and reports the\noutcome in that packet (section 4.3.3). One code is reachable only here -\n`0x97` for a publisher already holding `broker.qos2.max_inflight_per_client`\nunfinished exchanges - and it does not end the connection, because such a\nclient is inside a bound rather than misbehaving and its next exchange\nsucceeds as soon as one of its own completes.\n\n**`0x87` is what a derived dead-letter channel answers a client** - it\ntakes records from its queue's own dead-lettering and from nowhere else,\nand reading one is untouched (\"The dead-letter channel\" below).\n\nThe row's other cause - a client whose roles do not allow the publish -\nneeds an `acl_file`; without one every authenticated client may use every\nchannel, which is what a broker with no such file does.\n\n**A publish beginning `$share/` is refused because nothing could ever\nread it**: every subscription starting with those characters is a shared\nsubscription, not a topic, and no channel can claim them either. Left\naccepted it is a record stored nowhere, delivered to nobody, and answered\nSuccess. Only the exact lower-case spelling is refused - `$SHARE/x` is an\nordinary topic an ordinary filter can subscribe to, and publishing there\nis allowed.\n\n**Sagüin does not send `0x10 No matching subscribers`**, which MQTT 5\nmakes optional; a broadcast nobody hears is answered `0x00`, where\nmosquitto answers `0x10`. The cost is a producer's only signal that it\nmistyped a channel name: `event/orders`, one letter short of the `events`\nchannel, is an ordinary broadcast topic, acknowledged on every publish,\nand nothing on the wire distinguishes it from a deliberate broadcast.\n\n**Sagüin defines exactly five topics in the reserved space**, and a publish\nto anything else under `$saguin/` is the 0x90 above:\n\n| Topic | |\n|---|---|\n| `$saguin/consumer/\u003cchannel>/seek` | a consumer moves its own position |\n| `$saguin/queue/\u003cqueue>/response` | a worker acknowledges or returns a delivery |\n| `$saguin/kv/get` | a client reads the current value of one topic on a `latest` channel |\n| `$saguin/sessions/disconnect` | an operator hangs up the client whose id is the payload |\n| `$saguin/catalogue/\u003cchannel>` | a client asks what one channel is: its type, and the filter it claims |\n\n**And two topics a client may *subscribe* to**, which are the whole of what\nis delivered under `$saguin/`:\n\n| Topic | |\n|---|---|\n| `$saguin/queue/\u003cqueue>` | a worker consumes that queue: one job at a time, at QoS 1 |\n| `$saguin/consumer/\u003cchannel>/seek/reply` | a consumer reads the answer to its own seek when it set no Response Topic |\n\nPublishing to either is refused with everything else under the prefix - the\nfive above are the whole of what a client may publish there. **A queue's two\ntopics sit one level apart and are opposites**: work is delivered on\n`$saguin/queue/\u003cqueue>`, which may only be subscribed to, and answered on\n`$saguin/queue/\u003cqueue>/response`, which may only be published to. A filter\nof `$saguin/queue/\u003cqueue>/#` would straddle both, so it is refused with\nevery other spelling that is not a queue's exact form.\n\n**Everything else in the space is a topic a client publishes to**, so a\nsubscription to one of those is a promise that can never be kept and is\nrefused rather than granted and left silent. Nothing reaches this space by\nwildcard at all: MQTT already stops `#` matching a topic beginning with\n` RFC 0002 - Channels and configuration | Sagüin documentation , so only a client spelling one of these two out arrives here.\n\n**A subscription to the seek reply can only ever carry that client's own\nanswers**, and that is a property of how the reply is sent rather than a\nrule to enforce: Sagüin writes it to the seeking client's socket rather\nthan publishing it, so nothing reaches a subscriber that did not seek.\nWhich is what lets the topic be one name rather than one per client - a\nclient id in a topic would be a string a stranger chose, and then a rule\nabout who may subscribe to whose.\n\n**The first two name a channel that can answer them** - a seek an append\nchannel, a response a queue - and a name that matches nothing, or the\nwrong kind, is refused with a Reason String saying which. **The third\nnames no channel**: its payload is the topic to read, and a topic already\nresolves to a channel (RFC 0003 has the verb). **The fourth names none\neither**, a hang-up being an act on nobody's records (\"Hanging up a\nclient\" below). **The fifth names a channel of any kind** (\"Asking what a\nchannel is\" below).\n\n**The reason code judges the topic; the handler judges the request.** What\nis *inside* the message is not judged here, and each handler answers it its\nown way: RFC 0003 gives a seek a reply on the client's own Response Topic,\nand says of a response that it is ignored and logged and that a worker is\nnever told whether it was applied.\n\nThe `PUBACK` is the only answer that always arrives, and for a seek it\ncarries the refusal too: a client like `mosquitto_pub` that publishes and\nexits can read nothing else (RFC 0003).\n\n**A publish asking for a retained message is kept where there is\nsomewhere to keep it.** On a topic an `append` or `latest` channel claims,\nthe channel is that store and the flag adds nothing. On a `queue` the\nmessage is taken as ordinary work and the flag dropped -\n`saguin_queue_retain_ignored_total` counts those, nothing on the wire\nreporting a dropped flag. On a broadcast topic the value is kept in the\nretained store (\"Retained messages on a broadcast topic\" under\n*Configuration*). From a client whose roles deny `retained` the publish\nis refused with a `DISCONNECT` carrying `0x9A` (Retain not supported) -\nMQTT makes it a protocol error, and `0x9A` is no code a `PUBACK` may\ncarry. RFC 0003 \"Retained messages\" has the whole of it.\n\nThe consequence worth knowing before reaching for it: **nothing in the\nreserved space can be watched from the wire.** Somebody debugging a\nworker's acknowledgements watches the worker, not the response topic.\n\n**A topic that is malformed rather than merely unwanted never reaches this\ntable at all.** A wildcard in a `PUBLISH` topic is a protocol error under\nMQTT 5 rather than a refusable publish, and the answer is a **DISCONNECT\nwith 0x82 (Protocol Error)**; a publish into the `$SYS` tree the broker\nkeeps about itself is answered **`0x87` (Not authorized)**; and a topic\nwhose bytes are not valid UTF-8 fails to decode before any of this runs and\nthe connection closes with nothing sent. `mosquitto_pub` validates the topic\nitself and sends neither of the first two, so through that tool this\nbehaviour is invisible.\n\n**The broker asks the same questions of its own publishers.** A Will, an\ninbound bridge record and a queue delivery do not arrive on a socket, so\nthe same checks run on the way into a channel and a refusal is logged\ninstead - or, for a bridged record, held against the bridge's own\nacknowledgement (*Bridges*). A Will naming `events/+` would otherwise be\nstored in a channel where it stopped every conforming consumer for ever.\n\nA record the storage could not keep is refused rather than acknowledged:\na producer told its publish failed can retry or raise an alarm, and one\ntold it succeeded can do neither. The code is **0x83** because nothing\nthe client sent is wrong - MQTT 5 gives it for exactly that - where 0x80\ndeclines to say why, and every refusal here says why.\n\nPublishes above `max_message_size` never arrive. The limit is enforced\nagainst the declared packet length before any buffer is allocated -\nearlier than any hook - so no rule in this document applies to such a\npacket. **The bound is advertised**: the `CONNACK` carries Maximum Packet\nSize, so a conforming client never sends the packet, and one that does is\ndisconnected with **0x95 (Packet too large)**. The producer's own write\nstill fails part-way - the broker stops reading at the declared length -\nand the reason arrives on the same connection, so a producer that reads\nbefore it retries has it.\n\n## The dead-letter channel\n\nEvery `queue` named `\u003cname>` derives an `append` channel `\u003cname>__dlq`\nautomatically. It is not configured, cannot be configured, and shares the\nqueue's storage provider so that the move out of the queue and into it is\none transaction.\n\n**Its filter is derived too, and so is the topic each dead-lettered record\ntakes.** A record keeps the queue's name for its channel and gains one\nlevel in its topic: `__dlq`, inserted where the queue's filter carries its\n`#`, or appended where it carries none.\n\n```\nqueue filter dead-letter filter\niot/water/+/inspect iot/water/+/inspect/__dlq\niot/water/+/+ iot/water/+/+/__dlq\niot/water/+/inspect/# iot/water/+/inspect/__dlq/#\njobs/# (the default) jobs/__dlq/#\n\na record: iot/water/w-7/inspect → iot/water/w-7/inspect/__dlq\n```\n\nOne rule for every shape of filter, and the `+` levels keep what they\ncaptured because they are the same positions - one level is inserted and\nnothing else moves.\n\n**The rewrite is what makes a dead letter readable, not decoration.** A\nsubscriber finds a channel through the filters and nothing else, so a\nrecord still carrying the queue's own topic would belong to the queue and\nbe refused to everyone who is not a worker. Where the queue's filter ends\nin `#` the derived filter lies inside it, and the rule two sections up\nsettles which holds a topic without anything further being said: `__dlq`\nis spelled out where the queue has `#`, so the dead-letter channel is the\nmore exact of the two.\n\n`__dlq` is a reserved level in consequence. **No filter an operator writes\nmay carry a level equal to `__dlq`**, and that is refused at startup; every\nfilter that does is one the broker derived.\n\nIt is an ordinary `append` channel in every other respect. Consumers\nsubscribe to the topics above - `iot/water/+/inspect/__dlq`, or one\ndevice's, or `iot/#` and read dead letters among everything else - replay\nit, and hold independent positions. A publish *to* one of those topics is\nrefused `0x87`, because a dead-letter channel takes records from its queue\nand from nowhere else: a record put there directly has no queue record\nbehind it, so a redrive that reads `saguin-dlq-channel` and\n`saguin-dlq-offset` to put work back finds neither. The same rule refuses\na bridge rule naming one. Its seek topic is\n`$saguin/consumer/\u003cname>__dlq/seek`, by name, as every seek topic is.\n\n## Every reason code, in one place\n\nThe answer to a publish is under \"Publishing\" and a subscription's under\n\"What each channel type admits\"; what is left - the answers that arrive\nbefore a session exists or that end one - is collected here.\n\n**At `CONNECT`, in the `CONNACK`:**\n\n| Code | Refuses | To a 3.1.1 client |\n|---|---|---|\n| *unacceptable protocol version* | Any version below `broker.mqtt.min_protocol_version`, and MQTT 3.1 always - answered in the refused protocol's own vocabulary, the broker working rather than a fault (RFC 0001) | `0x01` |\n| `0x81` Malformed packet | A `CONNECT` with its reserved flag set or a Will QoS of 3, one whose Will QoS is set with no Will to govern, or one whose user name is not well-formed UTF-8 | `0x05` Not authorized; the user name `0x04`, being a broken credential |\n| `0x82` Protocol error | A `CONNECT` with Will Retain and no Will, or a 3.1.1 one carrying a password and no user name, which that protocol forbids | `0x05` Not authorized |\n| `0x83` Implementation specific error | A `CONNECT` carrying a `saguin-` User Property Sagüin does not read (RFC 0003 \"The reserved prefix on a client's own packets\"); or one the provider `broker.session.storage` names failed for, other than for room - a resumed session it could not read, a Will it could not write (\"Every session's state: `broker.session`\"), the ending of what a session begun new replaces (RFC 0003 \"Sessions\"), or what it still owes the client id from before - its last connection's disconnect (RFC 0003 \"Last Will\"), a record a connection that never completed wrote over, or what an ended session left in its channels (RFC 0003 \"Sessions\") - or the cursor of a shared group a resumed session holds and that has none | `0x03` Server unavailable, for the storage failure (3.1.1 has no User Properties to carry the other) |\n| `0x85` Client identifier not valid | A client id over `limits.max_client_id_length`, or a zero-length one asking for a session that outlives the connection - Clean Start 0, `cleanSession = 0` (MQTT-3.1.3-8) - because a session cannot be kept under no name, or one holding a control character (U+0001-U+001F, U+007F-U+009F), which mosquitto and EMQX refuse too | `0x02` Identifier rejected |\n| `0x86` Bad user name or password | A wrong credential and a missing one alike, so an unauthenticated caller cannot learn which user names exist; and a name holding U+0000 or a control character - a certificate's Common Name, a proxy's, or a `CONNECT` user name - which is nobody's (\"TLS on a listener\") | `0x04` Bad user name or password |\n| `0x87` Not authorized | A Will from a client whose roles deny `will` (*Taking a feature away*), or on a topic its roles do not allow it to publish to (*What a client may do*; RFC 0003 \"Last Will\") | `0x05` Not authorized |\n| `0x89` Server busy | `limits.max_connections` reached, on a tcp, tls or unix socket that found a place in the overflow budget and sent the opening bytes of a `CONNECT` (\"How long a socket may wait to send CONNECT\"); any other socket, and every ws one, is closed with nothing written | `0x03` Server unavailable |\n| `0x8C` Bad authentication method | A `CONNECT` naming an Authentication Method. Sagüin runs no enhanced authentication, so every method named is one it cannot continue, and a client library answered this falls back to an ordinary `CONNECT` | - (3.1.1 has no property to name one) |\n| `0x90` Topic name invalid | A Will that could never be delivered: the reserved `$saguin/` space, a dead-letter channel, over `max_topic_length` or deeper than `max_topic_levels`, or a topic name nothing may publish to - one holding `+` or `#`, or in the `$SYS` tree (RFC 0003 \"Last Will\") | `0x05` Not authorized |\n| `0x97` Quota exceeded | A Will the provider `broker.session.storage` names has no room for: MQTT can tell a client its session ends with its connection, and has no way to say \"your Will is not held\" (\"Every session's state: `broker.session`\"); or a resumed session holding a shared group that has no cursor, which that provider has no room to give it (RFC 0003 \"Broadcast\") | `0x05` Not authorized |\n| `0x9A` Retain not supported | A retained Will aimed at a broadcast topic from a client whose roles deny `retained` - where a retained publish is refused too (RFC 0003 \"Retained messages\") | `0x05` Not authorized |\n| `0x9B` QoS not supported | A QoS 2 Will from a client whose roles deny `qos2`, which would be told Maximum QoS 1 (MQTT-3.2.2-12) | `0x05` Not authorized |\n\n**A `CONNECT` that cannot be read is answered with the code the reading\nstopped on**, once it has said it speaks MQTT 5: `0x81` Malformed packet\nor `0x82` Protocol error - a property a `CONNECT` may not carry, a\nproperty sent twice or with a value its definition forbids, a length that\nruns past the packet, bytes left after the last field - and `0x95` Packet\ntoo large for one over the Maximum Packet Size. One refused before it says\nwhich version it speaks, or below MQTT 5, is closed with nothing written:\n3.1.1 has no return code for either.\n\n**3.1.1 has five return codes and Sagüin has more reasons than five**, so\n`0x05` is shared: the Will refusals and the malformed-flag refusals all\nland on it, because each is genuinely \"this connection may not do what it\nis asking to do\" - the closest of the five, rather than a code picked to\nfill the column. A `CONNACK` is all a 3.1.1 client gets: there is no\nReason String and no server `DISCONNECT` to follow it with.\n\n**A zero-byte client id is admitted on both protocols when it asks for no\nsession** - `cleanSession = 1`, Clean Start 1 - because a session ending\nwith its connection needs no name to be kept under. MQTT 5 then assigns\nan id and says so in the CONNACK's Assigned Client Identifier, a property\n3.1.1 does not have. Asking to *keep* a session under no name is the\n`0x85` row above, on either protocol.\n\n**`0x04` and `0x05` are kept apart on purpose**: a credential that does\nnot work and a Will the connection may not register are fixed by\ndifferent people in different files. The credential refusal keeps `0x04`;\neverything about what the connection may *do* is `0x05`.\n\n**A `DISCONNECT` from the broker**, which is the answer where no reply\npacket exists:\n\n| Code | Ends the connection because |\n|---|---|\n| `0x80` Unspecified error | A `PUBREC` for a QoS 2 delivery that the store could not record: no `PUBREL` is sent, and the exchange resumes with the session (RFC 0003 \"Broadcast\") |\n| `0x81` Malformed packet | A packet that cannot be read: a `PUBLISH` with both QoS bits set, flags MQTT reserves, a length that runs past the packet, bytes left after its last field, or a property the packet may not carry |\n| `0x82` Protocol error | A packet MQTT makes a Protocol Error as it is read, such as a `SUBSCRIBE` asking for Retain Handling 3, or any packet carrying a property twice that MQTT allows once or with a value its definition forbids (a Payload Format Indicator of 2, a Receive Maximum of 0). A publish whose topic name carries `+` or `#`, or one naming a topic alias this connection never registered - which every client that reconnects and keeps using its aliases does, since a new connection starts with none. A `SUBSCRIBE` carrying no filters at all: a zero-length filter is not a filter the broker dislikes but a field the client did not send, which MQTT calls a Protocol Error rather than `0x8F`'s business. A shared subscription asked with No Local, which MQTT-3.8.3-4 makes a Protocol Error (RFC 0003 \"No Local\"). An `AUTH` packet, no connection here having a method to re-authenticate with |\n| `0x83` Implementation specific error | A guard that should never fire: the substrate granted a SUBSCRIBE carrying a reserved property or a partition declaration Sagüin refused (RFC 0003), and the connection ends rather than deliver on a grant that lies. And a `PUBREL` whose release the store could not write: the message stays held (RFC 0003 \"Exactly once\"). And a session begun new whose ending of what it replaces the store refused after its `CONNACK` was written: the client's next `CONNECT` begins it again (RFC 0003 \"Sessions\") |\n| `0x8E` Session taken over | Another client connected with the same client id - enforced where the session lives, in the substrate |\n| `0x8F` Topic filter invalid | The same guard, for a shared subscription or a queue's form this connection cannot have, granted anyway |\n| `0x93` Receive Maximum exceeded | More QoS 2 exchanges unfinished than the Receive Maximum its `CONNACK` gave (MQTT-3.3.4-7); an exchange is unfinished until its `PUBREL` is answered |\n| `0x94` Topic alias invalid | An alias above the advertised maximum of sixteen, refused where the packet is parsed |\n| `0x95` Packet too large | A packet above the advertised Maximum Packet Size - and a bridge reads this one from its source, where it stops the link (\"The two brokers' `limits` blocks have to agree\"). And the other way: a channel's record, or a queue's job, over the Maximum Packet Size of the subscriber or worker it was for; a shared group's member it does not fit is passed over instead (RFC 0003 \"When a record is too large for a subscriber\") |\n| `0x97` Quota exceeded | A `PUBREL` whose release was refused for room: its channel filled after the `PUBREC`, or a retained broadcast's value finds no room in the retained store's provider. The message stays held, and the `PUBREL` sent again on the next connection completes it once there is room (RFC 0003 \"Exactly once\") |\n| `0x98` Administrative action | Somebody hung this client up: an operator holding `disconnect` published its client id to `$saguin/sessions/disconnect`. The session is untouched, so reconnecting resumes it |\n| `0x9A` Retain not supported | A retained publish to a broadcast topic, at either QoS, from a client whose roles deny `retained`: `0x9A` is not a code a `PUBACK` may carry, at QoS 0 there is no `PUBACK` at all, and acknowledging would claim a store the client may not use. A queue is not this case - there the flag is dropped and the work is taken |\n| `0x9B` QoS not supported | A QoS 2 publish from a client whose roles deny `qos2`, which its `CONNACK` told Maximum QoS 1 (MQTT-3.2.2-11). The same guard, for a queue subscription granted against Sagüin's refusal of its form or its QoS |\n| `0x9E` Shared Subscriptions not supported | A `SUBSCRIBE` containing a `$share/` filter from a client whose roles deny `share`, which its `CONNACK` told Shared Subscription Available 0 - a Protocol Error (MQTT 5 section 3.2.2.3.13). The same guard closes a connection the substrate granted one anyway |\n\n#### Every reason code Sagüin sends, by code\n\nThe four tables above are organised by packet, which is the right shape\nwhen you are reading about a feature. This one is the other question, and\nthe one somebody has at three in the morning: **I have a code, what could\nit be?** It is a reference rather than a second account - the causes below\nare the ones stated above, gathered under the number a client actually saw.\nA publish at QoS 2 reads its `PUBACK` rows off its `PUBREC` (\"Publishing\").\nEvery filter a `SUBACK` refuses is counted in\n`saguin_subscriptions_refused_total`, under the code decided for it - the\none below, even where a 3.1.1 client was sent `0x80` (RFC 0005).\n\n| Code | Sent on | What it means here |\n|---|---|---|\n| `0x00` | `CONNACK`, `SUBACK`, `PUBACK` | Accepted. On a `SUBACK` it is also the granted QoS |\n| `0x01` | `SUBACK` | Granted at QoS 1, and what a queue's form is always granted at |\n| `0x02` | `SUBACK` | Granted at QoS 2 - never on a queue's form |\n| `0x80` Unspecified error | `PUBREC`, `DISCONNECT` | On a `PUBREC`, an exactly-once publish from a connection whose client id another connection now holds. On a `DISCONNECT`, a `PUBREC` for a QoS 2 delivery the store could not record, so no `PUBREL` was sent |\n| `0x81` Malformed packet | `CONNACK`, `DISCONNECT` | A `CONNECT` with its reserved flag set or a Will QoS of 3, one whose Will QoS is set with no Will, or one whose user name is not well-formed UTF-8; any packet that cannot be read - a `PUBLISH` with both QoS bits set, flags MQTT reserves, a length that runs past the packet, bytes left after its last field, a property the packet may not carry |\n| `0x82` Protocol error | `CONNACK`, `DISCONNECT` | A packet MQTT makes a Protocol Error as it is read, such as Retain Handling 3, a property sent twice that MQTT allows once, or one with a value its definition forbids; a topic name carrying `+` or `#`, or an alias a connection never registered; a `SUBSCRIBE` carrying no filters at all; a shared subscription asked with No Local (MQTT-3.8.3-4); a `CONNECT` with Will Retain and no Will, or a 3.1.1 one with a password and no user name; an `AUTH` packet, which is answered on a `DISCONNECT` because no connection here has an authentication method to re-authenticate with |\n| `0x83` Implementation specific error | `CONNACK`, `SUBACK`, `UNSUBACK`, `PUBACK`, `DISCONNECT` | QoS 0 on a queue's canonical form; a `saguin-` property Sagüin does not read, on a `CONNECT` or a `SUBSCRIBE`; a partition declaration it refuses (RFC 0003); storage that could not keep the record, or at a `CONNECT` could not read the session being resumed, write the Will, end what a session begun new replaces, write what it still owes the client id from before, or give a shared group a resumed session holds the cursor it has none of; on an `UNSUBACK`, a session store that could not forget the filter, which stays subscribed (\"Every session's state: `broker.session`\"); on a `DISCONNECT`, an exactly-once release the store could not write, the message kept held (RFC 0003 \"Exactly once\"), or that ending refused after the `CONNACK` (RFC 0003 \"Sessions\"); a seek whose payload is not an offset or a time |\n| `0x85` Client identifier not valid | `CONNACK` | A client id over `limits.max_client_id_length`, a zero-length one sent with Clean Start 0, or one holding a control character |\n| `0x86` Bad user name or password | `CONNACK` | A wrong credential, a missing one where the listener requires it, and a name holding U+0000 or a control character |\n| `0x87` Not authorized | `SUBACK`, `PUBACK` | A filter or a channel the client's roles do not allow; a derived `__dlq`, which takes its records from the queue it failed in |\n| `0x89` Server busy | `CONNACK` | `limits.max_connections` reached and the socket found a place in the overflow budget and sent a `CONNECT`; otherwise it is closed with nothing written, always on a ws door |\n| `0x8B` Server shutting down | `DISCONNECT` | The broker is stopping |\n| `0x8C` Bad authentication method | `CONNACK` | A `CONNECT` naming an Authentication Method. Sagüin runs no enhanced authentication, so every method named is one it cannot continue, and a client library answered this falls back to an ordinary `CONNECT` |\n| `0x8E` Session taken over | `DISCONNECT` | Another client connected with the same client id |\n| `0x8F` Topic filter invalid | `SUBACK`, `DISCONNECT` | A string that is not a topic filter at all - `+` or `#` taking less than a whole level (MQTT-4.7.1-1, -2) - from an MQTT 5 client; a filter deeper than `limits.max_topic_levels`, on a `SUBACK`; a queue subscription that is not the canonical form, one asked with No Local (RFC 0003 \"No Local\"), or one this connection cannot have; `$share` in any spelling but MQTT's own lower case, or not shaped `$share/\u003cShareName>/\u003cfilter>` |\n| `0x90` Topic name invalid | `CONNACK`, `PUBACK` | Over `max_topic_length`; the reserved `$saguin/` space; a Will that could never be delivered |\n| `0x92` Packet Identifier not found | `PUBCOMP` | A `PUBREL` for an exactly-once publish nothing holds any more: it outlived `broker.qos2.expires_after`, its session ended or did not come back, its channel was removed, or it was released before a restart its `PUBCOMP` did not survive (RFC 0003 \"Exactly once\") |\n| `0x93` Receive Maximum exceeded | `DISCONNECT` | More QoS 2 exchanges unfinished than the Receive Maximum the `CONNACK` gave |\n| `0x94` Topic alias invalid | `DISCONNECT` | An alias above the advertised maximum |\n| `0x95` Packet too large | `CONNACK`, `DISCONNECT` | A packet above the advertised Maximum Packet Size, the `CONNECT` included; a channel's record or a queue's job over the Maximum Packet Size of the subscriber or worker it was for |\n| `0x97` Quota exceeded | `CONNACK`, `PUBACK`, `PUBREC`, `SUBACK`, `DISCONNECT` | On a `PUBACK`, **five different things** - see below - and on a `PUBREC` the same five, joined by a publisher already holding `broker.qos2.max_inflight_per_client` unfinished exchanges. On a `DISCONNECT`, an exactly-once release its channel has filled too far to take since the `PUBREC`, or a retained broadcast's release whose value the retained store's provider has no room for: the message stays held, and the `PUBREL` sent again on the next connection completes it once there is room (RFC 0003 \"Exactly once\"). On a `CONNACK`, a Will the session store has no room for, or the cursor of a shared group a resumed session holds and that has none. On a `SUBACK`, the session store `broker.session.storage` names is at its `max_bytes`, for every filter the packet would add or change (\"Every session's state: `broker.session`\"), or a filter that would take the client past `limits.max_subscriptions` (\"How many topic filters one client may hold\") |\n| `0x98` Administrative action | `DISCONNECT` | An operator hung this client up with `disconnect`; the session survives and reconnecting resumes it |\n| `0x99` Payload format invalid | `PUBACK` | Payload Format Indicator says UTF-8 and the payload is not |\n| `0x9A` Retain not supported | `CONNACK`, `DISCONNECT` | A retained Will, or a retained publish, aimed at a broadcast topic by a client whose roles deny `retained` |\n| `0x9B` QoS not supported | `CONNACK`, `DISCONNECT` | A QoS 2 Will or a QoS 2 publish from a client whose roles deny `qos2`; a queue subscription granted against Sagüin's refusal of its form or its QoS |\n| `0x9E` Shared Subscriptions not supported | `DISCONNECT` | A shared subscription from a client whose roles deny `share` |\n\n**A 3.1.1 client sees none of these numbers.** Its `CONNACK` carries one of\nthe five return codes mapped above, its `SUBACK` carries a granted QoS or\n`0x80` for a refusal, and its `PUBACK` carries nothing at all. That last\none is why a publish refusal it cannot be told about is a closed connection\nrather than an acknowledgement.\n\n**`0x97` is the one to know about.** MQTT 5 has no finer code, so Sagüin\nanswers all five of these with it: a channel at its `max_bytes`, its\nprovider at its `max_bytes`, a publish over `max_header_count`, one over\n`max_header_bytes`, and a client over its `publish_rate` or\n`publish_bytes` - and a QoS 2 publisher meets the same five on its\n`PUBREC`, joined by a sixth: it already holds\n`broker.qos2.max_inflight_per_client` unfinished exchanges. A publisher\ncannot tell them apart from the code alone - the Reason String beside it\nsays which, but MQTT lets a client ask not to receive one and Eclipse\nPaho asks not to by default.\n\n**An operator can tell them apart**, and that is deliberate rather than\nincidental: a metric is not bound by MQTT's code set, so\n`saguin_publish_refused_total` labels a rate refusal `publish rate exceeded`\nand the inflight refusal `inflight allowance exceeded`, where the rest are\n`quota exceeded`. \"Am I throttling my own devices\" and \"is a channel full\"\nare opposite problems with opposite fixes, and an operator watching refusals\nclimb should not have to guess between them. The same series is what makes a\nQoS 0 refusal visible at all, since nothing reaches the publisher there.\n\n**And what the `CONNACK` promises when it says yes**, so a client never\nhas to guess:\n\n**Every row below is an MQTT 5 `CONNACK` property.** A 3.1.1 `CONNACK`\ncarries a session-present flag and a return code and nothing else, so a\n3.1.1 client is told none of it. Every limit is still enforced against it -\nthat is what the refusals above are for - with one exception, Server Keep\nAlive, because a ceiling that works by telling the client cannot be\nenforced against a client that cannot be told. `limits.max_keepalive` says\nwhat happens instead.\n\n| Property | Value |\n|---|---|\n| Maximum QoS | **Absent**, which MQTT reads as 2 (section 3.2.2.3.4): the property carries only 0 or 1, so a broker offering 2 omits it. 1 for a client whose roles deny `qos2` (*Taking a feature away*) |\n| Topic Alias Maximum | 16, per connection (RFC 0001) |\n| Maximum Packet Size | `limits.max_message_size` |\n| Session Expiry Interval | Capped at `limits.max_session_expiry`, and a client that asked for more is told the granted value. 0 for a client whose roles deny `persistent` |\n| Server Keep Alive | Sent when `limits.max_keepalive` is set and the client asked for more, or for zero. A client must then use the value it is given (MQTT-3.1.2-21) |\n| Retain Available | 1 where some channel or the retained store can keep one, 0 where nothing on this broker can - and where it is 1, one topic's fate is still its channel's business, which is what the refusals above are for. A client whose roles deny `retained` is told the same, because that denial is about broadcast topics only |\n| Shared Subscription Available | 0 for a client whose roles deny `share`, and absent for every other, which MQTT reads as 1 |\n\n## Configuration\n\nOne file. Example:\n\n```yaml\nbroker:\n id: edge-1\n log_level: info\n\n mqtt:\n listen:\n tcp:\n address: 0.0.0.0:8883\n tls:\n cert_file: /etc/saguin/tls/cert.pem\n key_file: /etc/saguin/tls/key.pem\n ws:\n address: 0.0.0.0:8083\n unix:\n path: /run/saguin/saguin.sock\n mode: \"0660\"\n password_file: /etc/saguin/clients.passwd\n allow_anonymous: false\n min_protocol_version: \"3.1.1\"\n\n operations:\n listen:\n tcp:\n address: 127.0.0.1:9090\n unix:\n path: /run/saguin/operations.sock\n mode: \"0660\"\n min_scrape_interval: 60s\n\n storage:\n default: local\n default_retention_period: 3d\n default_retention_bytes: none\n providers:\n local:\n type: sqlite\n file_path: /var/lib/saguin/saguin.db\n volatile:\n type: memory\n max_bytes: 64MiB\n snapshot_dir: /var/lib/saguin/snapshots\n\n retained:\n storage: local\n retention_period: none\n\n limits:\n max_message_size: 1MiB\n max_topic_length: 1024\n max_client_id_length: 256\n max_header_count: 32\n max_header_bytes: 8KiB\n max_connections: 10000\n max_subscriptions: 10000\n max_topic_levels: 200\n max_session_expiry: 30d\n max_keepalive: none\n publish_rate: 200\n publish_bytes: 128KiB\n write_timeout: 5s\n connect_timeout: 10s\n max_connect_size: 100KiB\n max_connect_rate: 500\n session_queue_bytes: 1MiB\n\nchannels:\n\n events:\n type: append\n filter: iot/+/+/events/#\n storage: local\n retention_period: 7d\n retention_bytes: 20GiB\n\n device-state:\n type: latest\n filter: iot/+/+/{state,location}\n storage: volatile\n retention_period: 30d\n\n presence:\n # No filter, so this channel claims presence/#.\n type: latest\n storage: volatile\n retention_period: none\n\n jobs:\n type: queue\n filter: iot/+/+/work\n storage: local\n max_bytes: 128MiB\n visibility_timeout: 30s\n job_expires_after: 6h\n retry:\n max_attempts: 5\n backoff: exponential\n backoff_base: 2s\n dlq_retention_period: 30d\n dlq_retention_bytes: 1GiB\n```\n\nThere is no `type: broadcast`. A topic no filter matches is broadcast, and\nconfiguring that would only be a way to get it wrong.\n\n`filter` is the one key above that decides what a channel *holds* rather\nthan how it holds it, and everything in \"The topic space\" applies to it.\nOf the four channels here, three divide one `iot/\u003csite>/\u003cdevice>/`\nhierarchy, which is the thing a channel name could not do: an event log,\ncurrent state, a work queue. `presence` claims a tree of its own, and\n`iot/hq/dev-1/diagnostics` is left as ordinary broadcast because nothing\nclaims it.\n\n### What this broker is called: `broker.id`\n\n**There is no default and it is refused when absent.** The id names this\nbroker in its own log lines and on `saguin_build_info`'s `broker_id` label,\nand a generated one would be worse than none: the first time two brokers'\noutput is read side by side, the question is which machine said a line, and\nan invented name answers it with something nobody can look up.\n\n### How much the broker says: `broker.log_level`\n\n`debug`, `info`, `warn` or `error`. Absent is `info`.\n\n```yaml\nbroker:\n log_level: debug\n```\n\n| Level | What it adds |\n|---|---|\n| `error` | Only what stopped working |\n| `warn` | Refusals a client was given, a bridge that cannot reach its peer, **and the line that takes each of those back** |\n| `info` | Startup, listeners, retention sweeps, connects and disconnects - **never a line per record** |\n| `debug` | Every bridge connection attempt, the detail behind a retry, and a line per record: each job offered, leased, acknowledged or returned, and each value deleted |\n\n**A line that retracts another is written at the level of the line it\nretracts** - an operator owed the moment a problem began is owed the\nmoment it ended, or a log read a week later reports an outage that ended\nin its first minute. `bridge link up again` is therefore a `warn` like\nevery other clear, while a first connection stays at `info` because it\nretracts nothing.\n\n**Some lines exist only at `debug`**: a bridge writes one line per failed\nconnection attempt there, and a bridge that cannot reach its peer at all\nis reported once at `warn` - the fault visible either way, the attempts\nbehind it not.\n\n**Nothing logs once per record at `info`**, because on a busy queue that\nis gigabytes a day on an edge box's disk: a queue handing out a hundred\njobs a second writes 1.9 GB a day of them being offered and acknowledged.\nThe counters carry those numbers (`saguin_queue_delivered_total`,\n`_acknowledged_total`, `_returned_total`; RFC 0005), and what an operator\nacts on - a dead-letter, an acknowledgement for a superseded delivery - is\na `warn`.\n\nAnything a value is refused for is named:\n\n```console\n$ saguin --check-config saguin.yaml\nconfiguration is invalid:\nsaguin.yaml:\n - broker.log_level \"verbose\": it is one of debug, info, warn or error; omit the key for info\n```\n\n**`SIGUSR1` re-reads this key, the `ws` listener's `same_origin` and\n`allowed_origins`, the credential files, and the TLS certificates and\nclient CAs.** Getting `debug` out of a misbehaving link otherwise costs a\nrestart, and a restart on a broker holding durable sessions costs every\nconnection and every record in flight - which is the same reason\nwithdrawing one device's password is on this signal too, under \"Who may\nconnect\". The broker logs the level it moved from and the level it moved\nto, and says in the same breath that addresses and storage are\nstartup-only - so an operator who edited two keys and signalled is told\nwhich one took. A file that does not parse changes nothing and does not\nstop the broker, because the likeliest reason it does not parse is that\nsomebody is still typing.\n\n`SIGHUP` is caught and does nothing - unhandled, its default disposition\nwould terminate the process with no shutdown and no snapshot - and it\nsays so at `warn`, because Mosquitto reloads on `HUP`, a closing terminal\nsends one, and a silent no-op reads exactly like a reload that worked.\n\n### Where the process writes its id: `broker.pid_file`\n\nAn absolute path, or absent for nowhere.\n\n```yaml\nbroker:\n pid_file: /run/saguin/saguin.pid\n```\n\nAbsent is the ordinary answer, and the demo ships it commented out for that\nreason: systemd, Docker and runit each track the process themselves and want\nno file. It is for the supervisor that reads one - an init script, a\n`kill -USR1 $(cat …)` to raise the log level or to re-read a credential\nfile, a monitor that only knows a number.\n\nThe path is absolute, refused if it is not, for the reason every other path\nhere is: two copies started from two directories would each write a different\nfile and each believe it held the only one.\n\n**If the file cannot be written the broker exits** rather than running\nunmanaged - something reads a pid file an operator asked for, and a stale\npid is a pid that is now somebody else. It is written after the signal\nhandlers are installed, or a supervisor that signals immediately would\nkill the broker it just started, and it is removed on a clean shutdown.\n\n**There is no `log_file` beside it**, and the asymmetry is the point: a pid\nfile holds one number and is replaced, where a log accumulates - and\ninvariant 13 says everything that accumulates is bounded. Bounding a log\nmeans a size, a retention count, compression and deletion: logrotate,\nreimplemented inside a broker and worse. Sagüin writes to standard output and\nlets the init system bound it. RFC 0005 has the table of what does that\nwhere, and the entry worth reading before shipping is Docker's, which bounds\nnothing until somebody sets `max-size`.\n\n### How long a delivery may wait on a consumer: `limits.write_timeout`\n\nThe longest the broker will wait to hand a packet to a client. A client\nthat has not taken it by then is disconnected, and the packet is undone\nrather than half-sent.\n\n**It governs every write to a client**, by two routes. Sagüin writes an\n`append` or `latest` delivery, a dead-letter record and every control\nreply. The substrate writes a queue's offers and every broadcast, and is\ngiven this same duration to arm on its own writes - which it does inside\nthe client lock it holds across them.\n\n**That second route matters more than it sounds.** The substrate holds a\nclient's lock across the write, and a write that never completes holds it\nfor as long as that client stays connected. Unbounded, that write freezes\nthe loop that offers every queue's records, expires jobs and runs both\nretention sweeps - this deadline is the bound.\n\nA queue worker disconnected this way has its leases returned as any\ndisconnected worker does; what it does *not* get is a lease clock, for the\nreason invariant 7 gives.\n\n```yaml\nbroker:\n limits:\n write_timeout: 5s\n```\n\n**The failure it prevents is a duplicate on a durable channel** - a\npublisher held to a stranger's socket past its own patience gives up\ninside the window and re-sends, whatever the window is set to. So a\npublisher is not one of the goroutines that can be held: it waits on\nstorage and on nothing another client does, lengthened only by\nthe commit in progress, the next one, and where set\n`publish_commit_interval` below, the operator's own figure. Invariant 16\ncarries the mechanism and the measurement.\n\n**Disconnecting is the answer rather than dropping**, because a timed-out\nwrite has already put part of a packet on the wire and there is no way to\ncontinue that stream. It costs the consumer nothing: a durable consumer's\nposition advances only on its `PUBACK`, so a consumer disconnected here\nreconnects and resumes at the record it had reached. Nothing is lost and\nnothing is duplicated.\n\nThat is the consumer granted QoS 1. One granted QoS 0 sends no `PUBACK`,\nso its position advances on the write instead and what was on the wire\nwhen the link broke is behind it (RFC 0003). Hanging up costs that\nconsumer the records it had not read, which is the trade QoS 0 is.\n\n**Dropping is what a broker without a stored position has to do.**\nSagüin can hang up instead precisely because a channel consumer has a\nposition: the records are still in the channel, and a reconnect replays\nthem. The bound is the same; what it costs is not.\n\n**Broadcast is the exception to the buffering, and to nothing else.** A\ntopic no channel claims has no cursor to resume from, so a subscriber\nloses what could not be sent to it. Two counters, two different moments\n(RFC 0005): `saguin_deliveries_dropped_total` is the outbound queue\noverflowing, `saguin_deliveries_refused_total` the step before it, and\nwatching one without the other sees half of what a slow broadcast\nsubscriber costs.\n\n**The deadline still reaches it.** Dropping is what happens instead of\n*buffering*, never instead of `write_timeout`: a socket that takes no\nwrite for that long is disconnected whoever is on it and whatever the QoS.\nA deaf QoS 0 broadcast subscriber therefore gets both, in that order -\nshed for as long as its socket accepts writes, hung up on when it stops.\nReading either half as the whole rule sends an operator looking for the\nwrong symptom, and a short measurement shows only the first.\n\nA broadcast subscriber served at QoS 1 or 2 never reaches the drop: it is\nheld within its session's bound instead (`limits.session_queue_bytes`\nbelow).\n\n**Why it is configuration.** It is the time a *slow link* is allowed, and\nSagüin's deployments are the ones where slow is normal - a ship, a\nsubstation, a vehicle on a metered connection. The default of `5s` is\ngenerous for any link that is working: a write only waits at all once the\nkernel's own buffer is full, which means the far end has stopped reading\nrather than merely being slow. `none` removes the ceiling and restores the\nbehaviour above, which is a choice an operator may make and should have to\nwrite down.\n\n### How long a socket may wait to send CONNECT: `limits.connect_timeout`\n\nA connection is not a client until its `CONNECT` arrives. Until then there\nis no keepalive to bound it, so without a bound a peer that opens a socket\nand sends nothing holds a descriptor for as long as it likes.\n\n**It holds a `max_connections` slot from the moment it arrives**, as\nmosquitto, NATS and EMQX count connections, so everything a connection\ncosts before authentication is bounded by the same number as everything\nafter it. A peer holding slots with silent sockets holds them until this\nbound closes them.\n\n**Past the slots there is one overflow budget**: `max_connections` or 32\nplaces, whichever is fewer, shared by every socket that arrives with\nevery slot taken and has not been admitted. A socket that finds a place\nwaits for a slot for up to 50 ms; a slot given back goes to the socket\nthat has waited longest, and one arriving meanwhile queues behind it\nrather than taking the slot first. One whose wait ends with no slot keeps\nits place while the opening bytes of its `CONNECT` are read, to be\nanswered `0x89` (Server busy; 3.1.1 `0x03`) with no credential checked.\nEvery socket that is not admitted is closed within 100 ms of its arrival,\nor `connect_timeout` where that is sooner, its wait included. A socket\nthat arrives with the budget full is closed at once, with nothing read\nand nothing written. So the sockets the broker holds are at most\n`max_connections` plus the budget, besides the one a door has just\naccepted and is closing for want of a place, and a connection whose end\nis decided, whose socket outlives its slot as below; and a client that\nhung up and connects again at once finds the slot it gave back. A stopping\nbroker admits nothing more and ends every wait at once.\n\n**It gives the slot back the moment the broker decides the connection\nends**, not when the socket closes: before it writes the `CONNACK`\nrefusing a connection or the `DISCONNECT` ending one, and as it reads a\nclient's own `DISCONNECT` or close. A client turned away or disconnected\ncan connect again at once, as with mosquitto. From then on the\nconnection is ending: nothing more is read from it, and nothing is\nwritten to it but the `CONNACK` or `DISCONNECT` that ends it. The socket\ncloses after that, so it outlives its slot by what comes between:\nwriting that last packet, which `write_timeout` bounds, and what must\nhappen before the close - storing what a client's `DISCONNECT` changed,\nor publishing the Will of one that dropped.\n\n**On a `ws` door a connection arrives as a socket before it is an MQTT\nconnection**, so its slot is taken when the socket is accepted, before the\nHTTP upgrade. One that gets no slot is closed with nothing written, since\nthere is no MQTT connection yet to answer; mosquitto and EMQX close it\nthe same way. The door speaks HTTP/1.1 only. A request the door\nanswers instead of upgrading - one that is not an upgrade, from a page\nwhose origin is refused, or malformed - gives its slot back before that\nanswer is written. **This bound runs once, from the moment the\nsocket is accepted**, over its TLS handshake, its HTTP request - any body\nthe request declares included - and the `CONNECT` that follows, together\nrather than each in turn: a socket that has not sent its `CONNECT`\n`connect_timeout` after it arrived is closed, whichever of them it is\nstill in.\n\n```yaml\nbroker:\n limits:\n connect_timeout: 10s\n```\n\nA connection whose `CONNECT` has not arrived within `connect_timeout` is\nclosed without a reply. Once it arrives the client's own keepalive governs\nthe connection instead. The default is `10s`.\n\nIt is a whole number of seconds, `1s` or longer. `none` is refused: this\nbound is the only thing between an idle socket and the broker's\ndescriptors.\n\n**A stopping broker does not wait for it.** Shutdown closes every\nconnection that has not sent `CONNECT` rather than waiting out this bound,\nso the snapshot written after the listeners close is not delayed by a\npeer that says nothing.\n\n### The largest CONNECT before authentication: `limits.max_connect_size`\n\nA `CONNECT` is read before its client has authenticated, so its size is\nmemory a stranger chooses. `max_connect_size` bounds the whole packet, fixed\nheader included:\n\n```yaml\nbroker:\n limits:\n max_connect_size: 100KiB\n```\n\nThe limit is checked against the length the `CONNECT` declares, before its\nbody is read. A larger one is answered `0x95` (Packet too large) and closed;\nan MQTT 3.1.1 client, which has no such code, is closed with nothing written.\nEither way the rest of the packet is not read.\n\nThe default is `100KiB`, or `max_message_size` when that is smaller, and the\nkey may not be set larger than `max_message_size`. A `CONNECT` whose fields\nstay within MQTT's own limits - a Will payload of up to 65,535 bytes, a\npassword of the same - fits. A fleet sending larger ones raises it; a broker\nshort of memory is better served by a lower `max_connections`, because much\nof what a handshake holds does not depend on the size of its `CONNECT`.\n\n### How fast connections are accepted: `limits.max_connect_rate`\n\n`max_connect_rate` is how many connections each MQTT listener accepts a\nsecond, with bursts of up to twice that, or `none` for no limit:\n\n```yaml\nbroker:\n limits:\n max_connect_rate: 500\n```\n\n**A connection over the rate waits; it is not refused.** It stays in the\noperating system's accept queue until the listener takes it, so a fleet that\nreconnects all at once after an outage is admitted more slowly rather than\nturned away. What it bounds is the work done for connections that have not\nyet authenticated, a password check each; memory is bounded by\n`max_connections` and `max_connect_size`.\n\nThe default is `500`. At that rate 10,000 devices reconnecting together are\nall admitted within about 20 seconds. A broker on a small processor lowers\nit; one expecting large reconnect storms on ample hardware raises it or\nwrites `none`. The operations listener is not limited by it.\n\n### How long a client may go quiet: `limits.max_keepalive`\n\nA client chooses its own keepalive, and MQTT lets a server answer with one\nof its own that the client must then use (MQTT-3.1.2-21). `max_keepalive`\nis that answer, as a duration, or `none` - the default - for no ceiling.\n\n```yaml\nbroker:\n limits:\n max_keepalive: 5m\n```\n\nA client asking for longer is given `5m` in its CONNACK, and so is a client\nasking for **zero**, which MQTT defines as \"never disconnect me for being\nidle\". Nothing is refused: the client is told the value and uses it.\n\nA client already inside the ceiling is left alone and its CONNACK carries no\nServer Keep Alive.\n\n**A 3.1.1 client keeps whatever keepalive it asked for** - Server Keep\nAlive is an MQTT 5 property, and a ceiling enforced against a client that\ncannot be told it is a flap loop written into the configuration file - so\nin a mixed fleet this ceiling is a property of the MQTT 5 half.\n\n**What the ceiling is worth is how quickly a dead client is noticed.** A\nsession the broker cannot tell is gone lasts until `max_session_expiry` - a\nmuch longer clock, and the one that releases a durable consumer's stored\nposition. A shorter keepalive makes the broker's own answer the answer.\n\nThe ceiling is a whole number of seconds and MQTT carries it in two bytes,\nso anything above 65535 seconds is refused rather than truncated:\n\n```console\n$ saguin --check-config saguin.yaml\nconfiguration is invalid:\nsaguin.yaml:\n - limits.max_keepalive \"20h\" is longer than MQTT can express, which is 65535 seconds or about 18 hours; write `none` for no ceiling\n```\n\n### How much a session may hold: `limits.session_queue_bytes`\n\nThe most one session may hold of deliveries its client has not acknowledged,\nas a size. Absent is `1MiB`, or `max_message_size` where that is larger.\n\n```yaml\nbroker:\n limits:\n session_queue_bytes: 4MiB\n```\n\n**What it counts is what a session is owed**: every QoS 1 and QoS 2\ndelivery on the wire and unacknowledged, every one waiting for room in the\nclient's Receive Maximum, and every one waiting while the client is away -\nfor broadcast to a session that outlives its connection, the messages\nafter its cursor in the broadcast log that its filters matched when they\nwere published (RFC 0003 \"Broadcast\"). Each is counted as its payload,\ntopic and properties, and beside them as the memory of its entry, measured:\nabout 600 bytes for one in the client's in-flight table, and 300 once for\na table that holds any; 80 for one owed from the broadcast log and waiting,\nits entry on the session's list; and for one owed from the log and on the\nwire, both of those and the 192 bytes that record its packet identifier,\nand 600 once while the session has any of those on the wire. So what holding\nthem costs is inside the bound however small the messages: at 128-byte\nmessages a 1MiB bound is about 4,000 waiting for a session that is away, and\nabout 500 on the wire for one that is connected.\n\n**Full means the oldest goes, and the publisher is never refused.** Past the\nbound a session gives up the oldest it is owed that is not on the wire - for\nbroadcast from the log, its cursor moves past it - so a device that comes\nback is sent the most recent of what was published while it was gone. Each\nis counted as `session_queue_full`. The publisher is answered once its\nmessage is written, and told nothing of what any session gave up, because\nnothing about its publish failed (RFC 0003 \"Broadcast\").\n\n**Some deliveries are never taken back.** A QoS 1 delivery on the wire to a\nconnected client stays until it is acknowledged: dropping it frees its packet\nidentifier while the client may still acknowledge it, and that\nacknowledgement would complete another message. A QoS 2 delivery stays for\nthe same reason with a worse outcome, since a client holding an identifier\nfrom an unfinished exchange discards a new message under it as a duplicate.\nAnd a channel's own record or a queue's offer stays, because Sagüin tracks\neach under its identifier - a position or a lease would be left pointing at\nnothing.\n\n**So no more than half the bound is written and unacknowledged.** A delivery\npast that half waits to be written, and is written as acknowledgements make\nroom, earliest first; one is always written when nothing is on the wire,\nhowever large. What waits can be taken back, so a client that stops reading,\nor reads and never acknowledges, holds its bound as the oldest it was written\nand the newest it was not, and is sent the newest once it recovers.\n\n**It is not disconnected for holding its bound.** Its memory is bounded, and\n`limits.write_timeout` hangs it up only once its socket stops accepting\nwrites - which, with a bound the kernel's socket buffers can absorb, it may\nnever do. It stays connected until it recovers or its session ends.\n\n**A session at its bound holding nothing it can give up is refused the next\ndelivery** - deliveries at QoS 2 - and each refusal is counted beside the\ndrops. **A channel's own record is not refused: it waits where it is.** An\nappend consumer's position holds and a `latest` value stays pending, past\nhalf the bound as any delivery is, and each is written as acknowledgements\nmake room - so a consumer that reads and never acknowledges holds half its\nbound of a channel's records, and loses none of them. A session holding\nnothing it cannot give up always has room, so a message as large as the\nbound is never refused for its size alone. Otherwise a session passes its\nbound only by the delivery being queued, and only until the oldest it can\ngive up is gone. There is no count of messages as well: a client's own\nReceive Maximum is the window, up to the 65,535 MQTT allows.\n\n**Sizing it.** The bound is memory the broker allocates for a session, and\nthe process's resident size runs about twice what it allocates, because the\nruntime keeps room to collect in: measured at up to 1.3MB per session at the\n1MiB default, with a thousand sessions each holding its bound, on both\nsession providers. A broadcast message is stored once in the log however\nmany sessions are owed it, and each session is charged its size and the\n80 bytes above, so at 1MiB a session that is away holds 64 messages of 16KB\nagainst the 4,000 of 128 bytes above. A deployment that broadcasts large\nmessages to subscribers that fall behind sizes the bound from the burst they\nmust ride out: half a second of 16KB messages at 500 a second is 4MiB.\n\n**`none` is refused, and so is a size below `max_message_size`.** The first\nis a session that stops reading holding everything it is sent, which is the\nunbounded buffer invariant 13 refuses; the second could never queue the\nlargest message a client may publish.\n\n`saguin_session_deliveries_dropped_total{cause=\"session_queue_full\"}` counts\nwhat was dropped, and `saguin_session_queue_messages` and\n`saguin_session_queue_bytes` what every session holds (RFC 0005).\n\n### How many topic filters one client may hold: `limits.max_subscriptions`\n\nThe most topic filters one client may hold at once. Absent is `10000`.\n\n```yaml\nbroker:\n limits:\n max_subscriptions: 100\n```\n\n**A filter is memory that no other bound counts**: measured at about\n2.3KB across the indexes that hold it. Without a count, one connection\nsubscribing to 100,000 filters held 201MiB. At the default, one client\nholds at most about 23MB of them.\n\n**The default is high on purpose.** Home Assistant subscribes once per\nentity, so thousands of filters on one client is ordinary. An edge box\nwhose devices each hold a few filters can set `100`.\n\n**What counts is what the client would hold after the SUBSCRIBE.** A filter\nit already holds is replaced and adds nothing. A filter named twice in one\npacket is one. A shared subscription is one filter like any other. A filter\ngiven up by `UNSUBSCRIBE`, or by a refusal for any other reason, makes\nroom.\n\n**Past it, each filter over the count is refused `0x97` (Quota exceeded)**,\nwith a Reason String naming the key, and the rest of the packet is answered\nas it would be. The connection stays up. A 3.1.1 client is sent `0x80`,\nthat protocol's one failure code. The refusal is said at `WARN` once per\nepisode: the first refusal since the client last added a filter. A\nre-subscribe to a filter it holds adds none.\n\nA value below `1` is refused, since it would refuse every SUBSCRIBE.\n\n**A lowered value bounds what is added, not what is held.** A session\nrestored holding more filters than a lowered `max_subscriptions` keeps them\nall, since ending it would drop what it is owed; its next SUBSCRIBE of a\nnew filter is refused until `UNSUBSCRIBE` brings it under. A lowered\n`max_topic_levels` is the same.\n\n### How deep a topic may be: `limits.max_topic_levels`\n\nThe most levels a topic name or a topic filter may have. A level is what\nlies between two `/`, so a string has one more level than it has `/`.\nAbsent is `200`. A 201-level topic a bridge brings in from a peer with a\nlaxer limit is dropped as a record Sagüin would never accept\n(`saguin_bridge_unstored_total{cause=\"never_accepted\"}`, RFC 0005).\n\n```yaml\nbroker:\n limits:\n max_topic_levels: 32\n```\n\n**Every level of a filter is memory.** The topic index holds a node for\neach, measured at about 600 bytes. `max_topic_length` does not bound\nfilters, and a SUBSCRIBE packet may carry a filter of 65,535 bytes. Without\nthis key, one filter of 30,000 levels took 17MiB and two seconds of CPU,\nand a client could hold `max_subscriptions` of them. At the default, one\nfilter holds at most about 120KB.\n\n**One rule, wherever a topic or a filter enters.**\n\n- A publish over it is refused `0x90` (Topic Name invalid), as one over\n `max_topic_length` is. That includes a record a bridge brings in.\n- A Will over it is refused at `CONNECT` with `0x90`.\n- A `SUBSCRIBE` filter over it is refused `0x8F` (Topic Filter invalid),\n or `0x80` on 3.1.1, and the rest of the packet is answered as it would\n be.\n- A point read's key over it is refused `0x90`.\n\nA shared subscription is measured by the filter after `$share/\u003cShareName>/`.\n\n**A dead letter is one level deeper than its record**, because its topic\ngains `__dlq` (\"The dead-letter channel\"), so a 200-level record on a queue\ndead-letters under a 201-level topic. The move writes it directly, not\nthrough a publish, so it is kept. A subscriber reaches it with a shallower\nfilter ending in `#`, such as `jobs/__dlq/#`. A point read of it is refused,\nand a Sagüin peer an `out` rule ships it to refuses it `0x90`, which the\nrule counts as `peer_refused` and skips.\n\n**The files are held to it at startup.** A channel's filter, a bridge\nrule's filter and `topic:` template, and a filter or topic in the\n`acl_file` deeper than this could match nothing a client may publish, so\n`--check-config` refuses the file and names the line.\n\n**A lowered value bounds what is added, not what is held**, as a lowered\n`max_subscriptions` does. A session restored holding a filter deeper than\nit keeps the filter, since ending the session would drop what it is owed,\nand an `UNSUBSCRIBE` is not held to it at all: it allocates nothing, and\nit is how a client gives up such a filter.\n\nA value below `1` is refused, since it would refuse every publish and\nevery SUBSCRIBE. There is no `none`: a topic tree deeper than 200 levels is\nnot one anybody writes, and the bound is what stops one client from taking\nthe broker's memory one level at a time.\n\n### How fast one client may publish: `publish_rate` and `publish_bytes`\n\nEvery other bound here is a **size**, and a size does not bound a\n**rate**: one client publishing as fast as it can is limited only by what\nit lands in. A channel at its `max_bytes` refuses with `0x97`, which is\nbackpressure and works - and **a broadcast topic has no size bound at\nall**, so a device with a firmware bug looping on one would be answered as\nfast as the broker could go, for as long as it liked.\n\nThat is invariant 13 arriving through the one thing no size key bounds -\nwork per second - and on the hardware Sagüin is for it is a likelier\nfailure than anything the size bounds catch.\n\n**Two units, because they bound different things and either alone leaves\nthe other hole open.** A thousand one-byte publishes cost a kilobyte and a\nthousand topic lookups, permission checks and store writes - which is what\nactually saturates Sagüin, and why its own benchmark is in messages a\nsecond. Ten one-megabyte publishes cost ten operations and ten megabytes. A\ndevice can exhaust the broker either way, so both are counted and whichever\nis reached first refuses. Either may be written alone.\n\n**Two places, and they compose.** The broker-wide figures are the floor\nevery client takes:\n\n```yaml\nlimits:\n publish_rate: 200 # messages a second\n publish_bytes: 128KiB # bytes a second\n```\n\nand the `acl_file` names the exceptions, beside the roles rather than in a\nblock of their own:\n\n```yaml\nusers:\n demo: [tour] # short form, unchanged\n \"device-*\":\n roles: [device]\n limits:\n publish_rate: 100\n publish_bytes: 64KiB\n \"gateway-*\":\n roles: [gateway]\n limits:\n publish_rate: 5000\n publish_bytes: 8MiB\n```\n\nOne pattern, one entry. A separate top-level block keyed by the same\npatterns would have every operator writing each device pattern twice - two\nplaces to keep in step, and a typo in the second one silently meaning no\nlimit at all.\n\n**The broker-wide figures apply to everybody**: anonymous clients, clients\nno pattern matches, and brokers with no `acl_file` at all - which is most of\nthem. **A named client replaces them, up or down.** Tightest-wins would\nmake the broker-wide figure a ceiling nobody could exceed, and a gateway\nthat legitimately publishes faster than a sensor is most of why per-client\nlimits exist.\n\n| The client | Takes |\n|---|---|\n| Anonymous, or no `acl_file` | the broker-wide figures |\n| Authenticated, no entry applies to it | the broker-wide figures |\n| The entry that applies carries `limits:` | that entry's figures, instead |\n| The entry that applies carries none | the broker-wide figures |\n\n**Which entry applies is settled once, under *Roles, and users matched by\npattern*,** and it is one rule for the whole entry rather than a rule of its\nown for limits: where several patterns match a client's user name, the one\nthat spells it out most exactly supplies both its roles and its limits, and\nthe others do nothing. The figures follow the user name and the budget the\nclient id, so a client taking over another's id publishes at its own\nfigures, not the ones the id's last holder had.\n\n**Absent means no bound, and that is the default.** There is no figure to\npick: what a deployment sustains is its storage's rate, and memory and\nsqlite differ by an order of magnitude. A default drawn from either would\nthrottle a legitimate fleet on the other. A written `0` is refused at\nstartup rather than read as \"unset\" - those mean opposite things, and one\nof them is a broker that refuses every publish.\n\n**A message larger than the whole per-second byte budget is admitted**, and\nspends it into deficit rather than being refused. Refusing it every time\nwould make `publish_bytes` a silent ban on large messages rather than a\nrate, which is a different feature and not this one; `max_message_size` is\nwhere a size ceiling belongs.\n\n**Over the rate is `0x97` (Quota exceeded), and the connection is left\nalone**: disconnecting turns a burst into a reconnect storm, and it is\nnot needed - the refusal happens before the topic is resolved, before any\nauthorization is asked and before anything is stored, so nothing\naccumulates whatever the client does.\n\n**A budget belongs to the client and survives its connection** -\notherwise a device could have as many second's-worth as it could open\nconnections, and a library that treats the `0x97` as fatal and reconnects\ndoes exactly that by itself. The budget is dropped once it has refilled,\nso a client that has genuinely been quiet loses nothing.\n\n**At QoS 0 there is no reply**, so such a publish is dropped and the\npublisher is told nothing - the same silence every refusal takes at that\nQoS, since the only packet a server could send unsolicited is the\n`DISCONNECT` this refusal exists to avoid.\n\nThe log line is written at debug rather than warn, because a client held to\nits rate produces one per refused publish and a limiter that floods the log\nunder load has moved the problem rather than solved it. The metric\n`saguin_publish_refused_total{reason=\"publish rate exceeded\"}` is where an\noperator sees it, and it is deliberately its own series - see the table\nbelow.\n\n### Channels split across files\n\nOne file stops being the right shape somewhere well below the ten thousand\nchannels a real fleet reaches. Channels belong to domains, those have\ndifferent owners and different rates of change, and a single file makes\nevery one of them a merge conflict. So `channels:` may also be written as\na list, whose entries are each a group of channels - a file named by\n`!include`, or channels written in place:\n\n```yaml\nchannels:\n - !include channels/telemetry.yaml\n - !include channels/vessel-07/jobs.yaml\n - events:\n type: append\n storage: local\n```\n\nAn included file holds the same thing the block above holds: channel names\nand their settings, and nothing else.\n\n```yaml\n# channels/telemetry.yaml\nreadings:\n type: append\n storage: local\n retention_period: 7d\n```\n\nThe single mapping is still the whole of it for a configuration that needs\none file, and nothing about it changes. The rules the list follows:\n\n- **Only under `channels`.** This is not a general preprocessor. A file\n that can include anything anywhere cannot be reasoned about, and the\n thing an operator wants to split is the list that grows.\n- **Paths resolve against the including file, never the process\n directory.** Otherwise `make demo` and a systemd unit disagree about\n what the same configuration means, which is the worst class of\n configuration defect.\n- **No nesting.** A master file names files, and those files name none -\n nesting buys little and costs cycle detection, a depth bound, and an\n error that explains a chain. Refused from both directions.\n- **A channel name defined twice is a hard error naming both files** -\n two domains each defining `events`, one silently winning, is the real\n hazard of the shape - and two channels carrying the same `filter` is\n the same hazard refused the same way.\n- **The order the files are named in decides nothing** - which channel\n holds a topic is the filters' own business (\"Which channel a topic\n belongs to\"), so a domain's file can be added, moved or sorted without\n moving anybody's records.\n- **No globs.** `!include channels/*.yaml` makes what the broker serves\n depend on a directory listing, and an editor's backup file becomes\n configuration. Explicit names cost one line each and can be reviewed.\n- **An unknown key stays a hard error**, and the message names the file it\n is in. Losing the position is most of what would make a split\n configuration worse than a single one.\n\nOne caveat on that last rule: an `!include`d file's positions are exact,\nwhere a group written in the *master* reports a line counted from the\nstart of that group, so a comment inside the group moves the answer.\n\n### Storage providers split across files\n\n`broker.storage.providers` takes the same two forms and follows the same\nrules, for the same reason: the team that owns a domain's channels usually\nowns the disk they sit on, and splitting the channels while leaving every\nprovider in one file leaves that team editing the master anyway.\n\n```yaml\nbroker:\n storage:\n default: local\n providers:\n - !include storage/local.yaml\n - !include storage/vessel-07.yaml\n```\n\nA channel in one file names a provider defined in another. Resolution is\nby name and that is the point of it - a provider is a name a channel\nrefers to, not a place in the document.\n\nOne rule is stronger here than for channels. **Two provider names on one\n`file_path` or one `snapshot_dir` are refused at startup, naming both\nproviders and both files** - two names on one store are two writers on\nit, each with its own bound and its own retention. Paths are compared\nafter cleaning, so `/var/lib/x` and `/var/lib/./x` are the same store; a\nsymlink can still make two paths one store, and what Sagüin refuses is\nthe case visible in the text.\n\n**One thing it cannot refuse: the total.** Two sub-operators each\ndeclaring 8GiB of memory on a 2GB box is not a clash - each provider is\nseparately reasonable and only the sum is not, and Sagüin does not know\nwhat the machine has. So the startup line states the total declared across\nmemory providers, in the one place somebody is looking, and the arithmetic\nbelongs to whoever owns the box.\n\n### How a sqlite provider commits publishes\n\nCommitting has a cost that does not depend on how much is in the\ntransaction, so paying it once per record is what bounds the write rate of\na provider.\n\n**When these two keys are absent, a `sqlite` provider collects without\nwaiting.** A publish that finds no transaction committing is stored at\nonce, in a transaction of its own. Publishes that arrive while one is\ncommitting are stored together in the next, which starts the moment it\nends. Nothing waits for company, so a lone publisher is answered as fast\nas with a transaction per publish, and under load a transaction holds\nwhatever arrived during one commit - the batch grows with the traffic\nrather than with a setting. One transaction holds at most 256 records;\nthe next arrival starts another, committed after it.\n\n`publish_commit_interval: none` gives every publish its own transaction.\nAn operator may instead ask for the publishes that arrive close together\nto be stored in one transaction, which waits for them:\n\n```yaml\n storage:\n providers:\n local:\n type: sqlite\n file_path: /var/lib/saguin/saguin.db\n publish_commit_interval: 2ms # absent to collect without waiting; `none` for a transaction per publish\n publish_commit_max_records: 32 # required with it, and there is no default\n```\n\n**A transaction closes at whichever of the two is reached first** - when\n`publish_commit_max_records` records have arrived, or when\n`publish_commit_interval` has passed since the transaction opened,\nwhichever happens sooner. Under load the record count is what closes it\nand the interval never fires; when little is being published the interval\nis what closes it, and it is the only thing that can, because a lone\npublisher never reaches the count.\n\n#### What each key accepts, and why it stops there\n\nEverything below is refused at startup, naming the key and the reason. A\nsetting whose wrong values are accepted in silence is a defect waiting for\nsomebody to write one.\n\n| | `publish_commit_interval` | `publish_commit_max_records` |\n|---|---|---|\n| **Absent** | collect without waiting - the default | must also be absent |\n| **Accepted** | `none`, or a duration from `1ms` to `1s` | a whole number from **2** to **`limits.max_connections`** |\n| **Refused** | anything longer than `1s`; anything unparseable; either key on a `memory` provider | `1` or less; more than `limits.max_connections`; absent while `publish_commit_interval` is set; set while `publish_commit_interval` is not |\n\n**Why `1s` is the ceiling on the interval**, and why single-figure\nmilliseconds is the useful range, is three things and the first surprises\npeople:\n\n1. **The wait is paid per connection, not per batch.** Sagüin stores one\n record per connection at a time - a connection's next `PUBLISH` is not\n read off the socket until its last is stored - so a client that\n publishes and waits for its acknowledgement gets exactly one record per\n interval, whatever else is happening. At `1s` that is one message a\n second from that client: twenty publishes issued at once down one\n connection against a `50ms` interval take 1.019s.\n2. **It is not a wait for readers, though it looks as though it should\n be.** A channel's reads run beside the writer (\"How a sqlite provider\n reads\", below), and even with `read_connections: 0`, where they share\n the one write connection and an open transaction does hold it, the\n leader waits out the interval holding *nothing* and opens the\n transaction only afterwards. So the connection is free for the whole\n interval and busy for the commit alone: with a publisher waiting out a\n `1s` interval, metrics scrapes on the same provider answer in under 3ms,\n and window reads on the collecting channel in at most 2.5ms with the\n wait doubled. What a reader does wait for is the transaction itself,\n which is longer with a batch in it - and that is bounded by\n `publish_commit_max_records`, not by this key.\n3. **Nothing beyond a few milliseconds buys throughput.** The record count\n is what collects a batch; the interval only decides how long a batch\n that will not fill waits before giving up.\n\n**Why `2` is the floor on the count.** A transaction that may hold one\nrecord still waits out `publish_commit_interval` before committing it -\nstrictly slower than a transaction per publish, while reading in a file as\nthough collecting were on. `none` is a transaction per publish, and leaving\nboth keys out collects without waiting.\n\n**Why `limits.max_connections` is the ceiling on the count**: a batch can\nnever hold more records than there are connections to have sent them, so\na larger count is unreachable by construction - every transaction would\nwait out the interval and commit a handful. The ceiling is not a target:\nwhat fills a batch is the connections publishing **at the same time**, so\na broker with `max_connections: 10000` and forty sensors wants a count\nnear forty.\n\n`publish_commit_interval` is what makes a transaction wait.\n`publish_commit_max_records` on its own is refused, because nothing would\nwait for the count and the file would read as though it set one.\n\n**And there is no default for `publish_commit_max_records`**, because the\nfigure that decides whether any of this helps is how many connections\npublish at the same time, and Sagüin does not know it.\n\nA batch that fills on the record count commits the moment the last record\narrives. A batch that does not fill waits out the interval and commits\nwhatever it has - so a count set above the traffic turns every commit into\na fixed wait and is **slower than not collecting at all**. Through the\nwire, by `BenchmarkPublishConcurrently`: each publisher a connection of\nits own, waiting for its acknowledgement before sending the next 128-byte\nrecord at QoS 1, the database on the machine's own disk - an AMD Ryzen 7\n260 on ext4 over NVMe, the broker and its clients held to 8 of its 16\nthreads - with the two right-hand columns at `publish_commit_interval:\n2ms`:\n\n| connections publishing | `none` | keys absent | `publish_commit_max_records: 256` | `publish_commit_max_records` at the publisher count |\n|---|---|---|---|---|\n| 1 | 7,100–10,100/s | 7,500–11,000/s | | |\n| 8 | 10,100–11,000/s | 25,000–36,600/s | 3,400/s - count above the traffic | 32,900–33,400/s |\n| 32 | 10,200–10,900/s | 56,000–69,500/s | 12,200/s - count above the traffic | 56,700–76,100/s |\n| 256 | 10,300–10,800/s | 89,300–118,200/s | 77,200–79,900/s - count met by the traffic | 73,600–92,200/s |\n\nRead the columns against `none`. A transaction per publish holds the\nprovider near ten thousand records a second however many connections\npublish. With the keys absent the batch is whatever arrived during the\ncommit before, so it grows with the traffic, and the rate with it - as\nfast as a count set to the publishers, without the deployment having to\nknow the number. Where a count is above the traffic it is never reached,\nso every transaction waits the full 2ms and commits a handful of records:\nat eight publishers that is **three times worse than `none`** and at\nthirty-two no better.\n\n**Set `publish_commit_max_records` at or below the number of client\nconnections you expect to be publishing at the same time.** Too low costs\na little - a batch closes on the count more often, and each is smaller -\nand too high costs everything.\n\n**Connections are what fill a batch, and in-flight depth is not** - one\nconnection is one record per interval whatever its Receive Maximum, and at\nQoS 0 alike, since the record is stored on the way through. The\n`mosquitto_pub -l` shape is the same trap from the other side: one\nconnection sending a file of lines is one record per interval, so it looks\nas though collecting made the broker slower. It did, for that client.\n\nTwo channels that want different answers go on different providers. Each\n`sqlite` provider is its own file, its own writer and its own setting,\nso a channel carrying sensor readings can collect and a channel carrying\ncontrol messages beside it need not.\n\nNothing about the durability changes. A record is stored before its\npublisher is told anything either way; the transaction is simply larger and\nthe wait for it is longer. A crash between the commit and the\nacknowledgements leaves records stored and unacknowledged, so the publisher\nre-sends and the channel holds them twice - which is the window a single\ncommit already has, widened by the interval.\n\nNeither key means anything on a `memory` provider, which has no\ntransaction to collect into, and both are refused there.\n\n### How a sqlite provider reads\n\nA provider writes on one connection. **A consumer's read of a channel runs\non a read-only connection beside it**, so a window of records never queues\nbehind a commit, and a commit never waits for a window being read.\n`read_connections` is how many of those a provider may open:\n\n```yaml\n storage:\n providers:\n local:\n type: sqlite\n file_path: /var/lib/saguin/saguin.db\n read_connections: 2 # absent for 2; 0 puts every read on the write connection\n```\n\n| | `read_connections` |\n|---|---|\n| **Absent** | `2` |\n| **Accepted** | a whole number from `0` to `8`. `0` opens none: every read runs on the write connection, behind the writer |\n| **Refused** | below `0`; above `8`; on a `memory` provider |\n\n**What it buys is consumers reading while publishers write.** On one\nconnection every consumer's read waits for the commit in progress, and\nevery commit for the reads queued ahead of it. Fifteen consumers draining a\nprovider that 16KB publishers keep saturated, on eight CPUs, at `0` and at\nthe default `2`:\n\n| | `read_connections: 0` | `read_connections: 2` |\n|---|---|---|\n| publishes a second | 1,291-1,374 | 1,940-3,013 |\n| deliveries a second | 19,337-20,587 | 28,712-43,822 |\n| `PUBACK` p50 | 94-96 ms | 42-62 ms |\n| process peak memory | 76-81 MiB | 130 MiB |\n| write-ahead log peak | 3 MiB | 111-315 MiB |\n\nAt sixty-four consumers `0` holds publishers to 359-375 a second, each\nwaiting a third of a second for its `PUBACK`, where `2` gives them\n1,078-2,230.\n\n**What it costs is memory and disk**, which is the reason to turn it down.\nEach read connection is a connection to the same database with a page\ncache of its own, opened when a read asks for it and kept for the next.\nAnd a reader holds the write-ahead log open while it reads, so under\nsustained reads the log grows beside the file until a checkpoint finds a\nmoment with no reader in it (RFC 0004). **A small box sets `1`**, which\nstill keeps reads off the writer, **or `0`**, one connection for\neverything. More read connections trade publishers for deliveries: past\none or two, each reader added takes CPU the writer needed, so publishes\nfall, while deliveries rise only where many consumers read at once. Above\neight a read waits for a CPU rather than for a connection, and each one is\nstill a page cache.\n\nWhat a read returns does not change with it. Each read is one snapshot,\nbegun for that read and ended with it: the channel's floor and the records\nabove it come from the same instant, and every read sees every commit\nbefore it, its own client's included. Only a channel's records are read\nthere. The broadcast log's reads decide what its deliveries acknowledge and\nremove, so they stay on the write connection, as does any read made inside\na write. RFC 0004 has what a reader does to the write-ahead log.\n\n### How often a sqlite provider forces its log to disk: `flush_interval`\n\nA provider forces its write-ahead log to disk on an interval, so that a\npower cut loses about one interval plus one fsync (RFC 0004 \"WAL, and\n`synchronous=NORMAL`\"):\n\n```yaml\n storage:\n providers:\n local:\n type: sqlite\n file_path: /var/lib/saguin/saguin.db\n flush_interval: 150ms # absent for 150ms\n```\n\n| | `flush_interval` |\n|---|---|\n| **Absent** | `150ms` |\n| **Accepted** | a duration from `10ms` to `1s` |\n| **Refused** | `none` and `0`, since there is no setting that turns it off; below `10ms`; above `1s`; anything unparseable; on a `memory` provider |\n\n**It trades acknowledged data lost to a power cut against disk wear.** A\nshorter interval loses less and writes more; RFC 0004 has both figures. An\nidle broker fsyncs nothing.\n\n### Which web pages may connect: `same_origin` and `allowed_origins`\n\nA browser opening a WebSocket sends an `Origin` header naming the site its\npage came from. Two settings on the `ws` listener decide which pages may\nconnect, and a page that fails is refused with `403` before the upgrade:\n\n```yaml\nbroker:\n mqtt:\n listen:\n ws:\n address: 0.0.0.0:8083\n same_origin: false\n allowed_origins:\n - https://dashboard.example.com\n - https://192.0.2.10:8443\n```\n\n| | |\n|---|---|\n| `same_origin` | The page's scheme, host and port must be the request's own: `https` on a TLS listener, and the host and port the browser connected to. Default `true` |\n| `allowed_origins` | The page's origin must be one of these. Default empty, which checks nothing |\n\n**When both are set, a page must pass both.** A dashboard served from\nanother site is therefore listed with `same_origin: false`. With\n`same_origin: false` and no list, every page is admitted.\n\nThese are the NATS server's `same_origin` and `allowed_origins`, with the\nsame meaning. **`same_origin` defaults to `true`**, so with no list only a\npage from the listener's own site is admitted.\n\n**Why a page from another site is refused.** Without the check, any site the\noperator's browser visits could open the listener as that browser: with the\nclient certificate the browser holds, or on a network where anonymous\nclients are admitted.\n\n**Behind a proxy that terminates TLS, `same_origin` refuses every page.**\nThe page is `https` and the request the listener receives is not, so the\nschemes never match. Such a listener is configured with\n`same_origin: false` and its pages listed.\n\n**`same_origin` alone does not stop DNS rebinding.** A page whose domain is\nre-pointed at the broker's address names that domain in both `Origin` and\n`Host`, so it is its own site. A list stops it, because the attacker's\ndomain is never on it.\n\n**A client that is not a browser is not affected.** paho.mqtt.golang,\nautopaho and MQTT.js under Node send no `Origin`, and a request without one\nis admitted. paho-mqtt for Python sends the broker's own scheme, host and\nport, which `same_origin` admits; with a list, that address is listed too.\n\n**An entry is an origin and nothing else**: `scheme://host` or\n`scheme://host:port`, a different port a different site, compared as a\nbrowser sends them. No wildcards, and an entry with a path, a query or a\nuser name is refused - a browser never sends one. `Origin: null` cannot\nbe listed: a sandboxed frame on any site sends it too.\n\nA refused upgrade is logged at `warn` with its origin, bounded. Both\nsettings exist on the `ws` listener only and are re-read on `SIGUSR1`; a\nfile that is not valid changes nothing.\n\n### TLS on a listener\n\nA `tcp` or `ws` listener may carry a `tls` block naming a certificate and\nits key, and so may the operations listener. Both paths are absolute, for\nthe reason every path here is: what a listener serves must not depend on\nwhere the broker was started from. Making the files is three openssl\ncommands per certificate, and the recipe closes the client-certificates\nsection below.\n\n```yaml\n tcp:\n address: 0.0.0.0:8883\n tls:\n cert_file: /etc/saguin/tls/cert.pem\n key_file: /etc/saguin/tls/key.pem\n```\n\n**Several listeners of a kind.** A kind with one door is written as a\nmapping under `tcp`, `ws` or `unix`, named after its kind. A kind with two\nor more is a list, and each entry then carries its own `name`, unique across\nevery MQTT door:\n\n```yaml\n tcp:\n - name: local\n address: 127.0.0.1:1883\n allow_anonymous: true\n - name: fleet\n address: 0.0.0.0:8883\n tls:\n cert_file: /etc/saguin/tls/cert.pem\n key_file: /etc/saguin/tls/key.pem\n```\n\nSo a plain port for local tools sits beside a TLS port for the fleet on\nthe same kind, each with its own address, TLS and auth - `password_file`,\n`allow_anonymous`, `tls` with `client_ca_file`, `proxy_protocol` on a\nUnix door, `ws`'s origins. `broker.mqtt.acl_file` stays one for the whole\nbroker, as does `limits.max_connections`, which is one count across every\ndoor of every kind. `broker.operations.listen` takes the same shape (RFC\n0005 \"The operations listener\").\n\n**The name is what a door is known by beyond its own block**: it is the\nlistener id `/v1/operations/users` reports a session under, what a log\nline's `listener=` names, and what a certificate error names -\n`broker.mqtt.listen.tcp` for the door a single mapping leaves named after\nits kind, `broker.mqtt.listen.tcp[fleet]` once a name says which of\nseveral. The single-mapping form is exactly this with the name \"tcp\",\n\"ws\" or \"unix\".\n\n`--check-config` refuses, by name with both doors quoted: a second door\nof a kind with none; two doors sharing one name - of any kind, including\na kind's own default name, so a `tcp` door named `ws` collides with a\nbare `ws:` map. **That rule is asked within one listener's doors at a\ntime, never across the two**: the MQTT doors (`tcp`, `ws`, `unix`\ntogether) are one closed set and the operations doors (`tcp`, `unix`)\nare another, because an operations door's name keys nothing an MQTT\ndoor's does - not `SetListenerCredentials`, not a route an MQTT client\nreaches, only the operations listener's own credential lookup - so an\nMQTT `tcp` door and an operations `tcp` door sharing a name are two\nunrelated things, not a collision.\n\nPorts and Unix paths are a physical fact rather than a naming one, so\nthose two checks *do* cross the two listeners: two TCP-speaking doors on\nthe same port, counting `tcp`, `ws` and the operations port together,\nwhere the hosts are equal or either is a wildcard (`0.0.0.0`, `::`, or\nno host at all - port `0` never clashes, since the kernel hands each\nlistener a different one); and two Unix doors, counting the operations\nsocket, at the same path - the kernel's port table and its filesystem\ndo not know which configuration block asked. Hosts are compared as addresses\n(`::ffff:127.0.0.1` is `127.0.0.1`) and paths as cleaned absolute paths\n(`/x/./r.sock` is `/x/r.sock`); a host name such as `localhost` is not\nresolved at check time, so it clashes only at the bind. A list of one door\nwith no name is named after its kind, as the mapping form is. A tcp or ws\naddress must read as `net.Listen` reads it (`host:port`, IPv6 in brackets, a\nport from 0 to 65535 or a service name `/etc/services` knows, `032010` and\n`+32010` being port 32010); ports are compared as numbers, a host is empty,\nan IP literal (a zone, if written, not empty; IPv6 multicast and zone-less\nlink-local refused) or an RFC 1123 host name, and a zone is ignored when\ncomparing hosts except on link-local addresses, where two zones are two\nsockets; a Unix path may not hold a NUL, and its last element must be a file\nname (not empty, `.` or `..`). **A plain start refuses exactly what\n`--check-config` does**, because both read the same configuration through\nthe same check - a broker never binds half its doors and leaves the clash\nfor the connections that land on whichever port happened to come up second.\n`SIGUSR1` re-reads every door's password file, certificate and client CA;\nwhich ports are open is startup-only.\n\n**Per listener rather than per broker.** The MQTT port faces a fleet whose\ncertificate is whatever their estate issues; the operations port faces a\nmonitoring system that is frequently somewhere else entirely. One\ncertificate for both would be one name for both, and a deployment with both\nwould have to choose which of them the name is wrong for. A Unix socket\ntakes none: the credential never leaves the machine, and a certificate there\nis ceremony rather than security.\n\n**There is no default and nothing is generated.** A self-signed certificate\nSagüin made for itself would be encryption no client can verify, so every\nclient would be told to skip the check - which looks secured, is not, and\nstops anybody asking again.\n\n**The certificate's key type decides most of a handshake's CPU.** Measured\nin Go on x86, a whole handshake, both ends in one process, costs 0.9 to\n1.1 ms with an RSA-2048 certificate and 0.24 to 0.43 ms with an ECDSA\nP-256 one (the lower figure of each is TLS 1.2, the higher TLS 1.3); the\nserver's signature alone costs 0.74 ms against 0.021 ms. An operator who\nexpects many clients to connect at once - a fleet reconnecting after a\nrestart, or a small ARM board - should prefer an ECDSA P-256 certificate.\n\n**The certificate is read at startup, and a bad one stops the broker**,\nwith an error naming the listener before a port opens - the alternative\nis a broker whose log says it is listening while clients fail somewhere\nelse, or one that carries on in plain text, the failure that looks like\nsuccess.\n\n**`SIGUSR1` re-reads it, and `client_ca_file` with it**, on every TLS\nlistener the operations one included, from the paths the broker started\nwith. A renewed certificate is served to every handshake after the signal;\na connection already open keeps the one it was handshaken with, so renewing\ncosts the fleet nothing - where a restart drops every connection, which is\nsix times a year for a certificate that renews every sixty days. A client\nCA replaced the same way refuses, from the next handshake, a certificate\nonly the old authorities trusted.\n\n**Nothing changes unless every listener's files load.** A renewal lands as\ntwo files written one after the other, so a signal between them meets a key\nthat does not match its certificate: the broker logs which listener and why,\nat `error`, and every listener goes on serving the pair it had. Each one\nthat is re-read logs its new certificate's expiry.\n\nA bridge's own `cert_file` and `key_file` are read at every handshake, so a\nrenewed pair is presented on the next reconnect with no signal at all.\n\n**`cert_file` holds the chain, not only the leaf**: a leaf issued by an\nintermediate is `cat leaf.pem intermediate.pem > cert.pem`, as nginx's\n`ssl_certificate` works. Getting it wrong fails on somebody else's\nmachine - \"unable to verify the first certificate\" at the client while\nthe broker's log says it is listening. The root itself is not included: a\nclient that does not already have it is not one this certificate can\nconvince.\n\n`min_version` is the lowest version a listener accepts, `1.2` or `1.3`,\nand 1.2 when it is not written. **It is a floor, not a pin** - a pin\nexcludes clients speaking something newer, the wrong way round for a fleet\nthat upgrades over years. There is nothing below 1.2 to choose.\n\n#### Client certificates\n\n`client_ca_file` names the authorities trusted when checking a certificate a\n**client** presents, which is what Mosquitto's `cafile` does. It is named\nfor that rather than copied across, because `ca_file` reads to most people\nas the chain a server sends, and here that is `cert_file`.\n\n**Absent, no client certificate is asked for or looked at** - an ordinary\nTLS listener. Present, mutual TLS, and `require_certificate` says whether\nevery client must have one:\n\n| | |\n|---|---|\n| no `client_ca_file` | encrypted, nobody's certificate examined |\n| `client_ca_file` | certificates verified; one is required |\n| `+ require_certificate: false` | verified if presented, and a client without one falls through to the password file |\n\nThat last row is the mixed mode a fleet migrating in batches needs, and it\nis the same shape as `allow_anonymous` beside `password_file`: presence is\nthe switch, and an explicit key relaxes it.\n\n**The certificate's name becomes the client's user name**, which is\nMosquitto's `use_identity_as_username`, and it is not optional here: a\ncertificate *is* a name, and a different one beside it would be two\nidentities for one client. The password file is not consulted for a\ncertificate client at all - the authority already made the statement that\nfile exists to make.\n\n**A certificate's name is its Common Name, or its first DNS name when it\nhas none**, taken exactly as the certificate states it - the rule the\noperations listener names a certificate by (RFC 0005), so one certificate\nis one identity at either door. CN is where an operator's own authority\nusually puts a name, but many authorities now issue a subject alternative\nname and nothing else. A verified certificate carrying neither names\nnobody: it authenticates nothing, falls through, and says so in the log.\n\n**A name holding U+0000 or a control character is nobody's** -\nU+0001-U+001F or U+007F-U+009F, from a certificate, a proxy or a `CONNECT`\nuser name alike - and the connection is refused `0x86`, with a warning\nsaying why: such a name could rewrite the log line recording it or pass for\nanother in one. It is the set mosquitto and EMQX refuse, and MQTT forbids\nU+0000 in a UTF-8 string outright [MQTT-1.5.4-2].\n\nA `require_certificate` written where no authority is named is refused at\nstartup: it reads as though certificates were being demanded, and nothing\nwould be.\n\n**Making the files is three openssl commands per certificate** (openssl 3;\n`-copy_extensions` is what carries the SAN into the signed certificate).\nOne authority signs both sides: `ca.pem` goes in the listener's\n`client_ca_file` and in every client's `--cafile`, and the same\ndevice-shaped pair serves a bridge's `cert_file` and `key_file` when two\nsaguins meet with mutual TLS:\n\n```sh\n# One authority for the estate, kept off the broker.\nopenssl req -x509 -newkey ec -pkeyopt ec_paramgen_curve:P-256 \\\n -days 3650 -nodes -subj \"/CN=my-estate-ca\" \\\n -keyout ca-key.pem -out ca.pem\n\n# The broker's certificate. The subjectAltName must say what clients\n# dial - a CN alone fails modern verification.\nopenssl req -newkey ec -pkeyopt ec_paramgen_curve:P-256 -nodes \\\n -subj \"/CN=broker\" \\\n -addext \"subjectAltName=DNS:broker.example.com,IP:192.0.2.10\" \\\n -keyout key.pem -out broker.csr\nopenssl x509 -req -in broker.csr -CA ca.pem -CAkey ca-key.pem \\\n -CAcreateserial -days 825 -copy_extensions copy -out cert.pem\n\n# A device's certificate. The CN is the client's name: it becomes the\n# user name and matches ACL patterns, and the password file is not\n# consulted for this client.\nopenssl req -newkey ec -pkeyopt ec_paramgen_curve:P-256 -nodes \\\n -subj \"/CN=device-7\" -keyout device-7-key.pem -out device-7.csr\nopenssl x509 -req -in device-7.csr -CA ca.pem -CAkey ca-key.pem \\\n -CAcreateserial -days 825 -out device-7.pem\n```\n\nand then a device connects with all three:\n\n```sh\nmosquitto_sub -V 5 -h broker.example.com -p 8883 --cafile ca.pem \\\n --cert device-7.pem --key device-7-key.pem -t 'state/#' -q 1\n```\n\nA plain client on the TLS port fails its handshake, a client offering no\ncertificate is dropped where `require_certificate` is on, the certified\none connects - and an `acl_file` rule on `%u` matches `device-7` from the\ncertificate alone, which is the Common Name doing exactly what this\nsection says.\n\n#### On a bridge, dialling out\n\nA bridge takes `ca_file`, the authority that signs its **peer's**\ncertificate, for a `tls://` or `wss://` peer a public root does not\nvouch for - which is what an estate running its own authority has. Absent,\nthe system roots are used.\n\n**There is no way to turn verification off**: a link that accepts any\ncertificate accepts whatever machine answered first, with the records\ngoing there. A `ca_file` beside an unencrypted peer is refused for the\nsame reason `require_certificate` without an authority is - it reads as a\nlink being verified when nothing is.\n\n`cert_file` and `key_file` are the other half: the certificate Sagüin\n**presents** when the peer asks for one, which is what Sagüin's own\nlistener asks for with `client_ca_file` set. Without them two saguins cannot\nbridge to each other under the strongest setting either of them offers. They\ngo together - a certificate with no key cannot be presented and a key with\nno certificate names nobody - and both are absolute, like every other path a\nconfiguration names. The device-shaped pair the recipe above makes\nserves here unchanged: a bridge is a client, and its certificate's Common\nName is its name at the far end.\n\nThe whole shape at once - a private authority checked, and an identity\npresented, which is mutual TLS with Sagüin as the client:\n\n```yaml\nbridges:\n head-office:\n peer: tls://mqtt.example.com:8883\n client_id: vessel-07\n\n # The authority that signs the FAR END's certificate, for an estate\n # running its own. Leave it out and the system roots are used.\n ca_file: /etc/saguin/tls/peer-ca.pem\n\n # The certificate saguin PRESENTS when the far end asks for one, and\n # its key. This is the half a listener's client_ca_file asks for.\n cert_file: /etc/saguin/tls/bridge.pem\n key_file: /etc/saguin/tls/bridge-key.pem\n\n topics:\n - filter: fleet/+/telemetry/#\n topic: telemetry/$1/$#\n direction: in\n```\n\n**A pair without a `ca_file` is allowed**, which is not the rule a listener\nfollows, and the difference is that here the two keys answer different\nquestions: the authority checks the far end, the pair identifies this one.\nA peer whose certificate a public root vouches for, reached with a\nprivate client certificate, is an ordinary arrangement, and refusing it\nwould be a rule with no failure behind it. A `cert_file` beside an\nunencrypted peer is refused for the same reason a `ca_file` is: there is\nno handshake for it to take part in, and it reads as though there were.\n\n**Sagüin presents what the configuration names, and lets the far end\nrefuse it** - handed to the TLS library as a candidate instead, a\ncertificate issued by the wrong authority is silently not sent, and the\nlink comes up carrying no identity wherever the peer takes one without.\n\n**A file Sagüin cannot read stops the broker**, exactly as a listener's\ncertificate does: a configuration error is answered before a listener\nopens, where a broker that started anyway would run a dead bridge for the\nlife of the process - the log saying up, the records not moving.\n\n### Which versions may connect: `broker.mqtt.min_protocol_version`\n\nThe oldest MQTT version this broker admits. Two values and no others:\n\n| Value | |\n|---|---|\n| `\"3.1.1\"` | the default. MQTT 3.1.1 and MQTT 5 clients are both admitted |\n| `\"5\"` | MQTT 5 only. A 3.1.1 CONNECT is refused with `0x01`, 3.1.1's own \"unacceptable protocol version\" |\n\n**MQTT 3.1 - protocol level 3 - is never admitted**, and no value admits\nit: its `SUBACK` has no failure code, so a queue subscription could only\nbe granted or dropped, and its `CONNACK` has no Session Present flag, the\nonly way to tell a returning consumer its position is gone.\n\n**It is written under `broker.mqtt`, once, for the whole broker.** Unlike\n`password_file` and `allow_anonymous` it cannot be written on a listener,\nbecause the version a device speaks is a property of the device rather\nthan of the door it arrives at: the same firmware reaching the same broker\nover TCP and over a Unix socket is the same decision both times.\n\n```yaml\nbroker:\n mqtt:\n min_protocol_version: \"5\" # the fleet is MQTT 5; turn 3.1.1 away\n```\n\nBoth values are accepted quoted or not - YAML reads `3.1.1` as text and\na bare `5` as an integer, and the loader takes both shapes.\n\n**What changes for a 3.1.1 client once it is in is not here.** RFC 0001\nsays what the two protocols each get; this key decides only whether one\ngets in at all.\n\n### Who may connect: authentication and the password file\n\n`broker.mqtt.password_file` names the clients that may connect, in\nMosquitto's format - the same file, hash for hash, so a deployment moving\nfrom Mosquitto keeps the credentials it already has. `saguin --passwd`\nmanages it; RFC 0005 has the operations half, which is a **different file**\nbecause an operator is not a device.\n\n| Key | |\n|---|---|\n| `password_file` | Absolute path, Mosquitto's format. Absent means there is nothing to authenticate against |\n| `allow_anonymous` | Whether a client offering no user name is admitted |\n\n**Both keys may also be written on a listener, where they win.** A broker\nwith one answer writes it once under `broker.mqtt`; a broker with two - a\nUnix socket whose file permissions guard it, beside a TCP port facing a\nfleet - writes the difference on the listener it belongs to. Those\npermissions are the socket's `mode`, `0660` if absent: owner and group,\nnobody else, the same key the operations socket takes.\n\n```yaml\nbroker:\n mqtt:\n password_file: /etc/saguin/clients.passwd\n allow_anonymous: false # the fleet must authenticate\n listen:\n tcp:\n address: 0.0.0.0:1883\n unix:\n path: /run/saguin/saguin.sock\n allow_anonymous: true # and this door is the file's to guard\n```\n\n**A listener naming its own `password_file` admits its own users and\nnobody else**, unless it also writes `allow_anonymous` - it does not\ninherit the broker's answer. **`allow_anonymous` takes its answer from\nthe password file when it is not written**: no file admits everybody, a\nfile admits only who it names. Written, it wins either way - `true`\nbeside a password file is the mixed mode a migrating fleet needs, where\nnamed clients authenticate and the rest are still let in.\n\n`false` with no password file is refused at startup. It is a broker nothing\ncan connect to, which is a configuration nobody means, and the alternative\nis learning it from every client being refused at once.\n\n**A wrong credential and a missing one are both `0x86`**, Bad User Name or\nPassword. Distinguishing them on the wire tells an unauthenticated caller\nwhich user names exist. The broker's own log says which of the two it was,\nbecause the operator holding the log is not the caller.\n\n**A password file is read at startup and re-read on `SIGUSR1`.** Adding a\nuser does not admit that user until the signal arrives: `saguin --passwd`\nwrites the file, and the broker answers from the copy it holds. So\nprovisioning a device in the field is an edit and a signal - never a\nrestart - and an operator who skips the signal watches a tool report\nsuccess while the device it added is refused.\n\n**The operations password file is re-read by the same signal, and so are\nthe certificates** (\"TLS on a listener\"). Which *file* a listener reads\nstays a restart: the signal re-reads the contents of the files the broker\nstarted with, and repointing `password_file` at another path is an\nordinary configuration change.\n\n**Nothing is applied unless every file loaded**, because a signal arrives in\nthe middle of an edit as often as after one. A file that cannot be read, one\nwhose hashes Sagüin does not understand, one that names nobody on a listener\nadmitting nobody else - any of these leaves the broker running exactly what\nit had, and says so at `error` naming the file.\n\nHashes are `$6 RFC 0002 - Channels and configuration | Sagüin documentation (SHA512) and `$7 RFC 0002 - Channels and configuration | Sagüin documentation (PBKDF2-SHA512) as Mosquitto writes\nthem. **argon2id is refused with an error naming the file and the line**,\nnot treated as a failed password: a migration from a file Sagüin cannot read\nmust not look like a fleet that has forgotten its passwords. **A user named\non two lines is refused the same way, naming both lines**, as Mosquitto\nrefuses it: only one of the two passwords could ever be the one checked,\nand `--passwd delete` removing that one would leave the other admitting the\ndevice it was run to withdraw.\n\n#### Behind a proxy that terminated TLS\n\n`broker.mqtt.listen.unix.proxy_protocol` reads a **PROXY protocol v2**\nheader from every connection: the client's real address, and the\ncertificate Common Name where the proxy verified one. Without it, a fleet\nbehind one proxy is a single local peer in every log line, every\n`max_connections` count and every authorization decision.\n\n**The name a proxy sends is an identity**, the way a certificate this\nbroker verified itself is one: it becomes the client's user name exactly\nas a Common Name does, so an `acl_file` rule about `%u` works whether TLS\nwas terminated here or in front. **The header carries a Common Name and no\nDNS name**, so a certificate the proxy verified that has no Common Name\nnames nobody here, even where Sagüin terminating TLS itself would name it\nby its first DNS name.\n\n| | |\n|---|---|\n| `proxy_protocol` | Read a v2 header. Unix listener only. Absent means none is expected |\n\n**This key exists on a Unix socket and nowhere else**: the header is the\npeer asserting who its client is, and a socket's file permissions already\ndecide who may assert it - the allowlist a TCP listener would need\nfirst.\n\n**With it set, a connection arriving without a header is refused.** The\nsocket exists because a proxy is in front of it, and serving one anyway\nwould mean the address in every log line depended on whether the proxy was\nworking.\n\n**A connection that has not sent its header within 5 seconds is closed**:\na proxy writes the header the moment it connects, so only a peer that is\nnot a proxy takes longer. This holds on the operations socket's\n`proxy_protocol` too, and the bound is fixed rather than configured.\n\n**v2 and not v1.** A v1 header carries addresses and no TLVs, so it cannot\nsay which certificate was verified - half of what this is for. A v1 header\nis refused naming the setting to change rather than reported as an absence,\nbecause `proxy_protocol on` is the natural thing to write and is v1.\n\n| | sends |\n|---|---|\n| HAProxy, `send-proxy-v2-ssl-cn` | addresses and the SSL TLV carrying the Common Name |\n| nginx 1.31.4, `proxy_protocol v2` | addresses, an authority TLV, the SSL TLV - version, Common Name, cipher, signature and key algorithm - and a CRC32C |\n| nginx, `proxy_protocol on` | a v1 header, which Sagüin refuses |\n\nnginx needs 1.31.4 or later for `v2` at all. Earlier versions have only\n`on`, and behind one of those Sagüin's answer is the refusal above.\n\n**What a correct pair looks like.** The proxy verifies the client and hands\nthe connection to the socket; the broker reads the header and takes the\nname. Four things have to line up, and each of them fails differently:\n\n```nginx\nstream {\n server {\n listen 8883 ssl;\n ssl_certificate /etc/saguin/tls/cert.pem;\n ssl_certificate_key /etc/saguin/tls/key.pem;\n ssl_client_certificate /etc/saguin/tls/clients-ca.pem;\n ssl_verify_client on; # 1\n\n proxy_pass unix:/run/saguin/saguin.sock;\n proxy_protocol v2; # 2\n proxy_timeout 20m; # 3\n }\n}\n```\n\n```yaml\nbroker:\n mqtt:\n listen:\n unix:\n path: /run/saguin/saguin.sock\n mode: \"0660\" # 4\n proxy_protocol: true\n```\n\n1. **`ssl_verify_client on`**, and this is the one that is not merely a\n misconfiguration. The proxy is the only thing that sees the certificate,\n so whatever it forwards is what Sagüin believes - the whole contract is\n that an authority checked the name and Sagüin is trusting the hop.\n\n nginx sends the Common Name whenever a client presented a certificate,\n **including one it did not verify**, and sends beside it whether it\n did. Sagüin believes the name only when the SSL TLV's client flags say\n the connection was TLS and a certificate was presented, and its verify\n result is `0`. With `ssl_verify_client optional_no_ca`, a certificate\n carrying `CN=device-7` and signed by an authority nginx does not trust\n arrives with verify result `21`: it names nobody, the connection is\n answered as one with no certificate - the password file, or anonymous\n where the listener allows it - and the broker logs `the proxy could not\n verify this client's certificate … verify_result=21`. The operations\n socket's `proxy_protocol` believes a name by the same rule.\n\n With `on`, the same certificate does not get past the handshake: nginx\n answers `client SSL certificate verify error: (21:unable to verify the\n first certificate)` and the broker never sees a connection. `optional`\n with a `ssl_client_certificate` that actually verifies is equivalent for\n clients that present one; `optional_no_ca` is the setting to refuse\n outright.\n2. **`v2`, never `on`.** `on` is v1 and is refused, naming the setting.\n3. **`proxy_timeout` above the largest keepalive in the fleet.** It\n defaults to ten minutes and severs an idle connection whatever keepalive\n the client negotiated; the broker sees a disconnect and a new session,\n and a fleet sees reconnect churn with nothing in its own logs to explain\n it.\n4. **The socket's permissions are the access control**, so the proxy's user\n must be in the group that owns it. Otherwise the proxy logs `connect()\n to unix:… failed (13: Permission denied)` and the broker logs nothing at\n all, because nothing reached it.\n\nWith a device holding `CN=device-7` and a role granting\n`topic: alerts/%u/#`, `alerts/device-7/fire` is answered `0x00`,\n`alerts/device-8/fire` and a channel the role does not name are answered\n`0x87`, and the broker logs `a proxy named this client …\nprincipal=device-7`. The rule is bound to the MQTT client, rather than\nenforced only at the proxy's door.\n\n**Certificate revocation stays the proxy's business**, because the proxy\nholds the handshake. Sagüin is told a name that was verified; it cannot\nre-check what it never saw.\n\n### What a client may do: authorization and the ACL file\n\nAuthentication says who a client is; this says what it may do. Without an\n`acl_file` every authenticated client may publish into every channel,\nsubscribe to every channel, drain a work queue, and hang any other client\nup.\n\n`broker.mqtt.acl_file` names an authorization file. It sits beside\n`password_file` because it answers the second half of the same question, and\nit is a **different file** from the password one so that the password file\nstays Mosquitto's format hash for hash.\n\n| Key | |\n|---|---|\n| `acl_file` | Absolute path. Absent means every authenticated client may do anything |\n\n**An `acl_file` requires something that authenticates a client, and is\nrefused at startup without one.** Authorization is a statement about an\nidentity, and with nothing to authenticate against there is no identity to\nmake it about - a file full of rules governing nobody reads as protection\nand is none.\n\nTwo things satisfy it, because Sagüin has two ways of naming a client: a\n`password_file`, or a listener with a `client_ca_file`, where the identity\nis the certificate's name. **A pure-certificate estate needs no\npassword file**, which matters because a certificate client is\ndeliberately not looked up in one (\"Client certificates\" above): requiring\none anyway would mean writing a file that authenticates nobody in order to\nhave any authorization at all, which is the very shape this refusal exists\nto prevent. For the same reason `allow_anonymous: true` beside an\n`acl_file` is refused, and the failure is worse than \"no rule applies\": an\nanonymous client's identity is the empty string, and a `*` pattern under\n`users:` matches it - so such a client is admitted and granted whatever\nthat wildcard grants, silently. With a password file and `allow_anonymous`\nunwritten an anonymous client is already refused, so this error fires only\nwhere an operator has said something contradictory out loud.\n\n**The same exposure reached without writing it is warned about, at the start\nand by `--check-config`**: a listener beside an `acl_file` that has no\npassword file and does not require a client certificate - a pure-certificate\nestate's Unix socket, a plain `ws` door, a `tcp` door with\n`require_certificate: false` - admits a client nothing identifies, whose\nidentity is the empty name and so gets whatever `*` grants. A user name such\na client types is not its identity: a name nothing checked is never one.\n\n#### Roles, and users matched by pattern\n\nRules are written into a **role**, and roles are given to **user names** -\nthe name a client authenticates under, matched by pattern. A fleet of ten\nthousand devices that are all the same kind of thing needs one rule, not ten\nthousand copies of it, and every copy is a chance to get one wrong.\n\n**The block is `users:` because its keys are user names** - never client\nids, the other string in the same packet, the one nothing proves, and the\none no rule here reads. A file that names the block `clients:` is refused\nat load, saying what to write instead.\n\n```yaml\nroles:\n telemetry-publisher:\n - channel: events\n filter: events/telemetry/%u/#\n allow: [write]\n job-worker:\n - channel: jobs\n allow: [consume]\n reader:\n - channel: events\n allow: [read, seek]\n - topic: alerts/#\n allow: [read]\n\nusers:\n \"vessel-*\": [telemetry-publisher, job-worker]\n dashboard: [reader]\n```\n\n`%u` stands for the client's identity: the password-file user name, or a\nclient certificate's name, which arrive in the same field.\n\n**`%c` stands for its client id**, and the two are not interchangeable.\n`%u` is proved - a password or a certificate stands behind it - so a rule\nresting on it rests on something. `%c` is whatever the device typed.\n\n**`%c` is for the deployment `%u` cannot serve**: a cohort sharing one\ncredential, where the proved name is the same for two hundred devices and\nthe only thing telling them apart is the id each chose. Scoping by it gives\neach device its own topics. It does not stop one taking another's, because\ndevices sharing a secret are one principal - what it costs the taker is\nthat MQTT hands the session over, so the device it displaced reconnects and\nsomebody sees it.\n\n**Written together they say two halves of one sentence**: the proved name\nfixes which cohort, and the chosen id picks the device inside it.\n\n```yaml\nroles:\n sensor:\n - topic: iot/%u/health/%c\n allow: [write, read]\n```\n\n`%c` alone would let any credential reach any device's topic. `%u` alone\ncannot tell two devices in one cohort apart. With `client_ids:` on the\nentry, the credential cannot even be carried by a device outside its\ncohort.\n\n**A name goes into a rule as one level, never as the rule's syntax.** Where\nthe name `%u` or `%c` would put in holds `+`, `#` or `/` - which a topic\nfilter reads as a wildcard or a level - that rule grants the client nothing,\nits other rules still apply, and the broker says so once, at connect;\n`saguin --acl` names the rule as withheld. A client id is whatever the device\ntyped, and `#` under `iot/health/%c` would otherwise be every device's topic,\nfor reading and for writing. mosquitto and EMQX refuse the same names. Each\nsubstitution is one pass, so a name holding `%c` is put in as written.\n\n**`saguin --acl` says what that scoping is worth**, wherever a rule uses\n`%c`, because a grant reading `iot/cohort-north/health/north-17` looks like\nper-device isolation and is not. It is separation by mistake rather than by\nforce, and only a file with `%c` in it is told so - a note printed on every\nfile is one nobody reads by the third time.\n\n**A deployment with one credential per device needs neither.** `%u` is\nalready the device, and `%c` adds nothing.\n\n**It is substituted wherever a rule names a filter - `topic:` and\n`filter:` alike - and never in a channel name.** A channel is declared once\nin `channels:` and cannot vary by client, so `%u` there would name a\nchannel per device, which is the fleet-sized copying roles exist to avoid.\nA rule whose `channel:` holds a `%` is a **startup failure** saying so,\nrather than the \"channel is not configured\" it would otherwise get, which\nis true and sends an operator to declare one.\n\n**A channel rule narrows with `filter:`, compared against the whole\ntopic** - the same comparison a `topic:` rule makes, and the same one a\nchannel's own filter makes. One mechanism, so a rule says what it grants on\nits own line rather than meaning something a reader has to work out from\nthe channel it names.\n\n**A filter matching none of its channel's topics is a startup failure**,\nnaming the rule's filter and the channel's. It grants nothing and reads as\nthough it granted part of the channel, which is the shape this document\nrefuses everywhere else. `%u` is compared as an ordinary spelled-out level,\nbecause the rule has to be decidable before any client connects.\n\n**A rule's filter is a well-formed MQTT topic filter**, checked at startup\nthe way a channel's is: `#` is only ever the last level, and `+` and `#`\neach take a whole level. A `#` in the middle is refused rather than read,\nbecause the matcher stands it in for everything that follows - so\n`alerts/#/page` would grant every `alerts` topic to a rule naming three,\nand read on the page as the narrowing it is not. A `topic:` rule is checked\nthe same way less the rule about the first level, which is a channel's\nalone: `topic: \"#\"` is how an operator says \"any broadcast topic\".\n\n**`{a,b}` works in a rule's filter and expands the way a channel's does**,\nso a rule may narrow a braced channel to one of its spellings or carry the\nchannel's own filter, and both mean what they look like. One braced rule is\none grant per spelling, which is what `--acl` prints.\n\n**It expands before `%u` and `%c` are substituted** - the other way\nround, a device calling itself `{s1,s9}` would turn one device's grant\ninto two. Expanding what the file says settles every grant before any\nclient is named.\n\n**Assignment by pattern is the half that saves the work.** With it,\nprovisioning a device is one line in the password file and no authorization\nedit at all - which matters here more than in most brokers, because neither\nfile means anything to a running broker until it is told to re-read them,\nand an avoided edit is an avoided signal.\n\n##### Which entry applies\n\n**Where several patterns match a client id, the one that spells it out\nmost exactly supplies the whole entry - roles and limits together - and\nevery other matching entry does nothing.** `device-7` beats `device-*`,\ncounted in literal characters, so `device-*` beats `*-7` for `device-7`.\n**Two patterns spelling out the same number and both matching one id are\nrefused at startup naming both** - which applies would otherwise be\ndecided by nothing an operator can read.\n\nOne entry, one rule: an operator reading a four-line entry can take it as\nthe whole of what that client gets.\n\n**The cost is that a narrow entry can take something away.** A file with a\nbroad `*` giving everyone a `base` role and a `device-*` beside it does\nnot give devices `base`; a `device-7` written to add one role drops the\nfleet's second; and a `device-7` carrying no `limits:` takes the fleet's\nbound off that device rather than inheriting it.\n\nThat is not a defect - it is the only way an exception can be expressed:\na grant-only model with no precedence cannot subtract. Every entry\ntherefore reads as the complete statement about the clients it names,\nwhich is why an entry with no `roles:` is still refused.\n\n**Nothing can refuse a shadowed entry at startup**, because shadowing is the\nintended behaviour rather than a mistake, and from inside the file a pattern\nthat matched and lost looks exactly like one that applied. So\n`saguin --acl \u003cconfig-file> \u003cuser-name>` names the entry in force and the\nentries it shadows, and says when a shadowed entry carried limits the\napplying one does not - the one form of this rule that loosens a bound\nrather than removing a permission, and the one that shows up as a device\nflooding the broker rather than as a device being refused.\n\n##### Which devices may carry a name\n\nEvery MQTT `CONNECT` carries two strings, and only one of them is proved:\n\n| | Chosen by | Proved by | What it names | Does authorization read it |\n|---|---|---|---|---|\n| Client Identifier | the client, freely | nothing | the **session**, so it can resume | never |\n| User Name | the client | its password, or the certificate carrying it | **who you are** | always |\n\nEverything above is about the second. A client id is whatever the device\ntyped, so a rule resting on one would rest on nothing.\n\n**`client_ids:` is where an operator ties the two together.**\n\n```yaml\nusers:\n \"cohort-north\":\n roles: [sensor]\n client_ids: \"north-*\"\n```\n\nThe name `cohort-north` may then be used only by a device whose client id\nmatches `north-*`. Anything else is refused at the door. One pattern or a\nlist of them, `*` standing for any run of characters as it does in the\nentry's own key, and `%u` substituted for the proved name - so\n`client_ids: \"%u\"` says \"connect under the name you logged in as\", which is\none line for a fleet of any size and is what makes a client id in a log\nline worth reading. A name holding `*`, the pattern's own wildcard, is never\nsubstituted: such a pattern admits no client id for that name.\n\n**Absent means any client id**, which is what most deployments want: a\ncredential per device already ties the two together and there is nothing\nfor this to add.\n\n**What it buys is the boundary between credentials, and that is the\nwhole of it**: a device holding the northern credential cannot connect as\na southern one. Inside a cohort it buys nothing and cannot - devices\nsharing a secret are one principal, and that impersonation stays noisy is\nalready the `%c` paragraphs' point.\n\n**The refusal is `0x86` (Bad user name or password)**, the same code a\nwrong password gets - \"your credential is good, your client id is wrong\"\nwould tell an unauthenticated caller the credential is good. The log line\ncarries the distinction, and that is where an operator looks.\n\n**A `client_ids:` that admits nobody is a startup failure** - an empty\nentry, or the key written with nothing under it. It reads as a restriction\nand is a locked door: every device holding the credential is refused at\nconnect, which looks like a broken fleet rather than a configuration\nmistake.\n\n**It can only refuse.** The user name decides which entry applies and what\nthat entry grants; this decides whether the connection happens at all. So a\nclient id, which nothing proves, never widens anything - it is asked after\nthe name is proved, and only ever subtracts.\n\n#### Three rule kinds, and only two of them are about topics\n\nEvery topic resolves to exactly one channel or to broadcast - filters\noverlap deliberately and the most exact of them holds a topic, which is\ninvariant 12's rule and is settled before any client connects. Two of the\nrule kinds hang off that partition, which is why they cannot conflict:\n\n- a **`channel:` rule** governs topics that resolve to that channel;\n- a **`topic:` rule** governs broadcast topics - the ones no channel claims.\n\nThe third is outside the partition because it is not about the topic tree\nat all:\n\n- a **`broker:` rule** governs a facility of the broker itself, and there is\n one of those today: `sessions`.\n\nNothing a `broker:` rule grants can be reached by any filter, and no\nfilter can confer it: **a role holding every verb on every channel and\n`topic: \"#\"` beside them still cannot hang anybody up**, which is the\npoint of a separate kind rather than a reserved topic somebody could\nmatch.\n\n**A `topic:` rule that lies wholly inside a channel is refused at startup,\nnaming both** - \"topic `iot/hq/+/events` is inside channel `events`, whose\nfilter is `iot/+/+/events/#`; write it as a channel rule\". That is\ninvariant 12's own shape one level up: refuse the ambiguity while the\noperator is looking, rather than resolving it consistently until something\nchanges underneath. `%u` is treated as an ordinary spelled-out level for\nthat check, since the rule has to be decidable before any client connects.\n\n**A rule that merely crosses a channel is not refused**, and the difference\nmatters more than it looks. `#` and `+/telemetry` reach every channel on\nthe broker and are still how an operator says \"any broadcast topic\". What\nthe refusal is for is the rule that could only ever have been about a\nchannel's own topics, which grants nothing and reads as though it granted\nthe channel. Nothing is left ambiguous by accepting the crossing form,\nbecause a topic is resolved before any rule is asked: a channel topic never\nreaches a topic rule, whatever that rule's filter covers.\n\n**A subscription is granted whole or refused** (*What a filter reaches*):\nits filter needs the grant for every channel it touches, and a filter that\nlies inside no one channel's filter needs a `topic:` rule covering it as\nwell, because it can match topics no channel claims. With the channels at\nthe top of this document, `iot/water/w-7/#` needs `read` on both water\nchannels and a `topic:` rule: `iot/water/w-7` itself is broadcast. A filter\nthat several channels together cover, and none alone, is judged the same\nway.\n\n#### The verbs\n\n| | verbs |\n|---|---|\n| `append` | `write`, `read`, `seek` |\n| `latest` | `write`, `read`, `delete` |\n| `queue` | `write` (submit work), `consume` |\n| `topic:` (broadcast) | `write`, `read` |\n| `broker: sessions` | `disconnect` |\n| `broker: features` | none: it takes `deny:` and nothing else (*Taking a feature away*) |\n\n**One pair everywhere.** `write` puts something in and `read` takes it out,\non every channel type and on broadcast alike, so a rule can be written and\nread without first looking up which type the channel it names happens to be.\n`seek` and `delete` are the two acts that are not a direction - moving a\nconsumer's position, and removing a key - and a client may hold either\nwithout the other.\n\n**`consume` is the one irregular word, and it is there to stop a\nmistake.** Reading an `append` channel costs nobody anything; taking a\njob holds the record for one worker. Spelled `read`, an operator granting\na dashboard `read` on a queue to watch the work would drain it instead -\ninvariant 11's failure through the ACL. **So `read` on a queue is refused\nat startup**, and the refusal says a queue cannot be watched at all - its\ndepth is a gauge on the metrics listener (RFC 0005) - because an operator\ntold only that `read` is not one of `write, consume` picks the nearer\nword, which is the grant the list was protecting them from.\n\n**The `$saguin/` verbs ride on the channel they name**, and get none of\ntheir own: a seek is `seek` on the append channel it moves, a point read\nis `read` on the latest channel it reads, and **a queue's response topic\nis part of `consume`** - separated, an operator could grant `consume` to\na worker that may never acknowledge, and every job it took would time\nout, retry and dead-letter.\n\n**`disconnect` is the one that rides on nothing**, and it is an exception\nbecause hanging up a client is not an act on anybody's records: there is no\nchannel to name, so there is no channel verb to hang it on. It is granted by\nthe `broker:` rule kind instead, which is why that kind exists at all - a\nverb with no subject would otherwise have to be spelled as a topic, and a\ntopic is something a wide filter can reach by accident.\n\n##### Hanging up a client: `broker: sessions`\n\n**A password file reaches a client at CONNECT and nowhere else**, so a\ncredential withdrawn and re-read is still a credential a connected device is\nusing: it was checked when the device arrived and is not checked again. The\ndevice goes on until something ends the connection, and a device with a good\nkeepalive has no reason to end it. **Hanging it up is the act that makes the\ndevice come back and be asked again** - and it is the only act in this\ndocument that is not about a topic.\n\n**On its own it withdraws nothing.** A hang-up costs the device one\nreconnection; what it may do when it returns is whatever the files say then.\nSo it is the second half of an act whose first half is editing those files\nand signalling the broker to re-read them, and the order matters: hung up\nfirst, the device is back with everything it had before the edit lands.\n\n```yaml\nroles:\n on-call:\n - broker: sessions\n allow: [disconnect]\n\nusers:\n oncall: [on-call]\n```\n\n**It is asked for by publishing to `$saguin/sessions/disconnect`, with the\nclient id to hang up as the payload.** In the payload rather than in the\ntopic, for the reason a point read puts its key there: a client id is\nwhatever the device typed - `/` and all - and a topic carrying one cannot\nbe parsed back into it.\n\n**The `PUBACK` says whether the request was taken**, as every reserved-space\nverb's does:\n\n| | |\n|---|---|\n| `0x00` | authorized, and acted on |\n| `0x83` (Implementation specific error) | no Response Topic to answer on; or the payload named no client id, or named the client that sent it |\n| `0x87` (Not authorized) | this client's roles do not carry `disconnect` |\n\n**What happened comes back on the Response Topic**, the way a seek's answer\ndoes - one word, on the topic the request named:\n\n| | |\n|---|---|\n| `hung-up` | it was connected; its DISCONNECT is on its way, and the connection ends within `limits.write_timeout` |\n| `no-such-client` | nothing is connected under that id |\n\n**Those two are worth telling apart**, which is the whole reason there is a\nreply at all: otherwise \"that device disconnected an hour ago\" and \"you have\nmisspelled the id\" are the same answer, and the second is the one an\noperator actually sends.\n\n**And they are told apart in a payload rather than by a reason code**:\n`0x10` would fit, and it is not dependable - a `PUBACK` may leave the\nreason-code byte out, so a code that means something can arrive as the\nabsence that means `0x00`. A payload cannot be omitted.\n\n**A client cannot hang itself up**: a connection ended by its own request\ncould never receive the acknowledgement, and every client can already end\nits own connection with the `DISCONNECT` its protocol defines.\n\n**At QoS 0 it works and is answered**, the answer being a reply rather\nthan a code; what QoS 0 cannot carry is a *refusal*, which is dropped in\nsilence - the trade the reserved space makes everywhere.\n\n**A request that names no Response Topic is refused**, `0x83`, and\nnothing is hung up - acting on it anyway would end somebody's connection\nand tell nobody it had happened.\n\nThe property is MQTT 5's, so **a 3.1.1 client cannot ask for this verb at\nall**: the refusal ends its connection - the rule every refusal a 3.1.1\nclient cannot be told about already takes - and the client hung up is the\none that asked, never the one it named. No fallback topic is kept here the\nway a seek keeps one, because no data depends on an operator's tool being\na 3.1.1 client.\n\n**It ends the connection and does not touch the session**: the device\nreconnects, resumes at its stored position, and loses nothing - the whole\ncost is one reconnection, which is what makes the verb safe to put in\nfront of an operator. Ending a session is a different act with a\ndifferent cost, and this document does not offer it.\n\n**An MQTT 5 client is told why**: a `DISCONNECT` carrying `0x98`\n(Administrative action). A 3.1.1 connection is closed with nothing on it,\n3.1.1 having no server-to-client `DISCONNECT`.\n\n**The broker's own in-process publisher cannot be hung up.** It is how the\nbroker publishes on its own behalf - a Will, a queue delivery, a bridge -\nrather than anybody's session, and a request naming it is answered\n`no-such-client` like any other id nothing is connected under.\n\n**Without an `acl_file` every authenticated client may do this** - a\nbroker with no authorization file has no privilege boundary, and this\nverb does not invent one to stand alone in.\n\n##### Taking a feature away: `broker: features`\n\n**MQTT gives every client five features a fleet's operator may not want\nevery device to have**: a session that outlives its connection, a Last\nWill, a shared subscription, a retained value on a broadcast topic, and\nexactly-once delivery. A `broker: features` rule takes any of them away\nwith `deny:`, each named for the broker block that configures it:\n\n```yaml\nroles:\n sensor:\n - channel: events\n allow: [write]\n - broker: features\n deny: [persistent, will, share, retained, qos2]\n```\n\n**A denied client is answered as a broker without that feature would\nanswer it**, in MQTT's own signal where there is one:\n\n| Denied | An MQTT 5 client | A 3.1.1 client |\n|---|---|---|\n| `persistent` | accepted, with a Session Expiry Interval of 0 in its `CONNACK`: its session ends with the connection | accepted, with its session made clean |\n| `will` | a `CONNECT` carrying a Will is refused `0x87` (Not authorized). One without a Will connects as usual | refused `0x05` |\n| `share` | told Shared Subscription Available 0 in its `CONNACK`. A `SUBSCRIBE` containing a `$share/` filter anyway is a Protocol Error, and the connection is closed `0x9E` (MQTT 5 section 3.2.2.3.13) | nothing changes: a 3.1.1 client is refused every `$share/` filter already |\n| `retained` | a retained publish to a broadcast topic is closed `0x9A`, and a `CONNECT` carrying a retained Will aimed at one is refused `0x9A` - what MQTT has a broker without retained messages answer. On a channel's topic nothing changes and Retain Available stays 1: the channel is the store, so the flag keeps nothing extra there, and a 0 would make every retained publish a Protocol Error on every topic | a retained broadcast publish closes the connection, and a retained Will aimed at broadcast is refused `0x05` |\n| `qos2` | told Maximum QoS 1 in its `CONNACK`. A `SUBSCRIBE` asking for QoS 2 is granted 1 and delivered at no more than 1; a QoS 2 publish anyway is closed `0x9B` (MQTT-3.2.2-11); a `CONNECT` carrying a QoS 2 Will is refused `0x9B` | a QoS 2 subscription is granted 1, a QoS 2 publish closes the connection, and a QoS 2 Will is refused `0x05` |\n\n**A feature is denied when any of the client's roles denies it**, and the\ndenial is broker-wide rather than per topic: a role written to keep a\nfleet's sessions short is not undone by another role that says nothing\nabout features.\n\n**It is asked once the client has authenticated** - asked of a name a\n`CONNECT` merely claims, the answer would tell an unauthenticated caller\nwhich names exist - so a wrong password is still `0x86`, whatever the\nname's roles deny.\n\n**A Will is refused rather than dropped**: quietly not arming an\nannouncement the device asked for is a half-grant nothing would report. A\npersistent session is not refused, MQTT having a way to say \"your session\nends with this connection\".\n\n**`deny:` is written only on a `broker: features` rule, and that rule\ntakes nothing else** - a feature is something every client has until a\nrole takes it away, so there is nothing for `allow:` to add. A denial\nanywhere else, an `allow:` here, a feature not among the five, or one\nwritten twice is refused at startup, `retain` by name because the feature\nis `retained`. `saguin --acl` and `/v1/operations/acl` show each denial\nbeside the rule it is written on.\n\n##### Asking what a channel is\n\n**A client addresses a channel by name everywhere except when it\npublishes or subscribes**, and only the configuration says which topics a\nname claims - so a library offering `publish(channel, key)` rather than\n`publish(topic)` has to learn the filter from somewhere that is neither\nthis file nor the operations listener, whose credential must not be a\ndevice's (RFC 0005). **`$saguin/catalogue/\u003cchannel>` is where it asks**,\nover the MQTT credential it already has:\n\n```\nPUBLISH $saguin/catalogue/readings\n Response Topic: where the answer goes - required\n Correlation Data: optional, echoed\n```\n\nThe answer is written to the asking connection, exactly as a point read's\nand a seek's are, so **no subscription is needed and no other client can\nreceive it**. It is JSON:\n\n| Field | |\n|---|---|\n| `name` | the channel's name, as asked for |\n| `type` | `append`, `latest` or `queue` |\n| `filter` | the filter this channel claims, **as written** - braces and all |\n| `verbs` | the verbs this client holds on it |\n| `pin` | a queue's subscription form, and only a queue has one |\n\n**The filter goes back as written rather than expanded.** `{device,sensor}`\nis a level with a fixed set of spellings, which is exactly what a caller\ncomposing a topic needs: a `+` takes any one level, a braced level takes one\nof these, and a trailing `#` takes the rest or nothing. The plain filters it\nexpands to are the broker's own business and would tell a caller less.\n\n**The answer is scoped to what the asking client may use, and the two kinds\nof no are the same no.** A channel this client holds no verb on and a\nchannel that does not exist are both an empty reply, so asking cannot be\nused to discover what is there. That is what makes this safe to answer over\na device's credential: a client already holding `write` on a channel learns\nthe topics it was already publishing to, and one holding nothing learns\nnothing at all.\n\n**A grant narrowed by `filter:` still learns the whole filter** - this is\na disclosure rule rather than a permission one, and the narrowing is\nenforced where it decides something: on the publish that leaves it.\n\n**And it is a permission answered as it stands, never a promise** - an\n`acl_file` can be re-read under a connected client. The filter itself\ncannot change under a live connection, channels being startup-only, so a\nclient may cache it for as long as the connection lasts.\n\n**The `PUBACK` carries every refusal**, for the reason a point read's does:\nthe reply carries an answer and nothing else, so an empty one can mean one\nthing.\n\n| | |\n|---|---|\n| `0x00` | answered - with the channel, or with the empty reply |\n| `0x83` (Implementation specific error) | no Response Topic to answer on |\n| `0x90` (Topic name invalid) | the name held a `/`, so it names no channel that could exist, and the request is not one |\n\n**Which means the request must be QoS 1**, since QoS 0 has no `PUBACK` to be\nrefused in; at QoS 0 it is dropped, as every other request in this space is.\nA 3.1.1 client cannot ask at all, the Response Topic being MQTT 5's.\n\n\n##### Withdrawing a device's access\n\nFour steps, in this order, and none of them restarts the broker:\n\n1. **Edit the two files.** `saguin --passwd delete \u003cfile> \u003cuser>` takes the\n credential away; removing the client's entry from the `acl_file`, or\n narrowing the role it names, takes the grants away. Either alone is\n worth doing and neither is complete without the signal below.\n\n2. **Signal the broker to re-read them**: `kill -USR1 $(cat\n /run/saguin/saguin.pid)`. The broker logs what it now holds - the paths,\n the number of users, the number of roles - or, if a file cannot be used,\n logs that at `error` and goes on running what it had. **Read the line\n before going on**: a signal that refused the file leaves the device with\n everything it started with.\n\n3. **Hang the device up**, by publishing its client id to\n `$saguin/sessions/disconnect`. The `acl_file` is already in force at this\n point - the device's next publish is refused whether or not it has\n noticed - but the password file is only read at CONNECT, so this is what\n makes the device connect again and be refused.\n\n4. **Read the reply.** `hung-up` means the device was connected and its\n DISCONNECT is on its way: the connection ends within\n `limits.write_timeout`, whether or not the device is still reading - the\n reply does not wait on a stranger's socket. `no-such-client` means\n nothing was connected under that id, which is either a device that had\n already gone or an id spelled wrong, and those are not the same thing to\n check next.\n\n**What each step alone leaves behind** is the whole reason to write them\ndown. Steps 1 and 2 without step 3 stop the device doing anything, but\nthe connection survives - authenticated on a credential that no longer\nexists - and a lease it holds goes back only when the connection ends.\nStep 3 without 1 and 2 costs the device one reconnection and changes\nnothing. A restart in place of step 2 works and ends every other\nconnection with it.\n\n**That same property is the verb's other use**, and it has nothing to do\nwith credentials: hanging up a worker that has wedged returns the jobs it was\nholding for another worker to take, rather than waiting out their visibility\ntimeout.\n\n**A device that authenticates with a certificate is withdrawn the same\nway, with one step missing and one limit.** Step 1 is the `acl_file` edit\nalone - there is no password-file entry to delete - and steps 2 and 3 are\nunchanged. **The limit is that it can still connect**: Sagüin has no\ncertificate revocation, so the handshake still succeeds and the device\ncomes back to no grants. Where the connection itself must be refused,\n`client_ids:` on the entry is what does it, and replacing the client CA\nis the same signal again (\"TLS on a listener\").\n\n#### Grant-only, for records\n\nDefault refuse; rules add; any match allows; a client's grants are the\nunion of the roles its entry names. A rule that can only grant cannot\ndisagree with another, so there is no precedence between rules to\nexplain - exactness decides one level up, and only which entry supplies\nthe roles.\n\n**The one denial is a feature**, on a `broker: features` rule, and it has\none rule of its own: any role denying it denies it, whatever the other\nroles say. It is not a refusal of an act on anybody's records - it takes\naway something MQTT gives every client - which is why it cannot collide\nwith a grant. *Taking a feature away* has the five there are.\n\n#### Authorization may only narrow\n\nThe substrate ORs its ACL hooks, and `OnACLCheck` carries Sagüin's\nstructural refusals - the queue form (invariant 4) and the wildcard\nboundary (invariant 11) - so a permission model able to widen them would\nput two workers on one job. **The structural rules are evaluated first\nand refuse regardless of any grant**: a rule may take permission away and\nmay never give it.\n\n#### What a refused client sees\n\nNothing new on the wire, and no code that did not already have a meaning:\n\n| | |\n|---|---|\n| a publish | `0x87` (Not authorized) on the `PUBACK` - the row the table under *Publishing* already carries |\n| a subscribe | `0x87` (Not authorized) in the `SUBACK`, per filter |\n| a Will | `0x87` (Not authorized) on the `CONNACK`, for a Will on a topic the client may not publish to. It is asked again, as that client, when the Will fires, because the file may have been reloaded since: refused then, it is logged and not published (RFC 0003 \"Last Will\") |\n| a QoS 0 publish | dropped: there is no acknowledgement to carry a refusal, the same trade the reserved space and `max_message_size` already make |\n\n**`0x87` on a subscribe says one thing**: this client, not this filter. A\nfilter of the wrong shape is answered `0x8F` or `0x83` instead, so\na developer is pointed at the filter rather than at their credentials.\n\n**The file is read at startup and re-read on `SIGUSR1`**, as the password\nfile is - and a re-read takes effect on connections already open, because\nthe file is asked on every publish and on **every delivery**, never only\nat CONNECT or SUBSCRIBE: a resumed session sends no SUBSCRIBE for the\nbroker to refuse. A subscriber refused mid-channel is stalled, not\nskipped, and the grant coming back is itself what wakes it; a queue job\noffered to a worker that may no longer take it goes back with its attempt\nunspent. Invariant 16 carries the mechanism and what each alternative\nloses.\n\n#### Explaining a decision: `--acl`\n\n```sh\nsaguin --acl /etc/saguin/saguin.yaml device-7\nsaguin --acl /etc/saguin/saguin.yaml cohort-north north-17\n```\n\nprints the roles that matched, the rules they carry, and the effective\ngrants. **The first argument after the configuration is the user name** -\nwhat a client authenticates as, and what every rule is written about - and\nnever the client id, which is a different string in the same packet and\nwhich no rule reads.\n\n**A client id is given only where a rule uses `%c`**, and then it must be:\nwithout one the substitution cannot be made, so the literal is left\nstanding and the grants printed below it match nothing. The command says\nthat in a sentence rather than printing lines that look resolved - a rule\nwritten on purpose, a configuration reporting ok, and a fleet refused\n`0x87` is the failure that would otherwise hide. It takes the **broker\nconfiguration** rather than the acl_file: that is where the `acl_file` is\nnamed and where the channels its rules are about are defined, and a rule\nnaming a channel is only answerable against them. Indirection has one cost -\n\"why can device-7 not publish?\" is two lookups and a pattern match\ninstead of a line in a file - and this is the answer to it, built with the\nfeature rather than after it. It is the same instinct as `--check-config\n--output`: where behaviour is resolved from several places, the resolved\nform is printable.\n\n**A list of user names on standard input answers them a line at a time**,\nwhich is the case the paragraph form is worst at. Seven tab-separated\ncolumns under a header naming them: the user name asked about, the role that\ngranted it, whether the rule is about a channel or a broadcast topic, which\none, the verbs, and the two figures that user is held to per second -\nmessages and bytes.\n\n**The two limit columns repeat on every row of a user**, because the figures\nbelong to the user rather than to the grant. That is the same reason the\nfirst column is always the input: a line stands on its own however the\noutput is cut about, and \"which of these forty devices is bounded\ndifferently\" is answered with a `cut` rather than forty runs of the\nparagraph form.\n\n**They carry `no bound` in words where nothing bounds one**, never `0` -\nthe one figure the configuration refuses - and never `-`, which this\nformat defines to mean granted nothing. **A line may carry a client id\nafter the user name**, separated by a tab or a space, for a file whose\nrules use `%c`; a line with one field is a user name alone. A channel\nrule that narrows with `filter:` carries both in the fourth column -\n`events (iot/+/events/device-7)` - because the filter alone would not\nsay which channel, and the channel alone would not say which of its\ntopics. `%u` is already substituted, because the substituted\nform is what the broker compares against. Against the files this\nrepository ships, so that it can be run - `examples/acl.yaml`, which the\ndemo stages where `examples/saguin.yaml` names it:\n\n```sh\n$ printf 'demo\\ndevice-7\\nnobody\\n' | saguin --acl examples/saguin.yaml\n# user role kind subject verbs rate bytes\ndemo tour channel events write,read,seek no bound no bound\ndemo tour channel state write,read,delete no bound no bound\ndemo tour channel presence write,read,delete no bound no bound\ndemo tour channel jobs write,consume no bound no bound\ndemo tour channel jobs__dlq write,read,seek no bound no bound\ndemo tour topic # write,read no bound no bound\ndevice-7 device topic iot/+/health/device-7 write,read 200 65536\nnobody - - - - no bound no bound\n```\n\n**A role carrying six rules is six lines**, as `tour` is, and a client\nholding two roles is a line per rule of each. **`device-7` is granted by\nits own entry**, not by `device-*`: the spelled-out pattern supplies the\nentry, and `%u` resolves to the name it authenticated under - a client\ncertificate's Common Name here, that device having no password anywhere.\n`jobs__dlq` is the dead-letter channel a queue derives; nobody configured\nthat name. **The list form does not say which entry applied**: its\nquestion is which of forty devices may publish, and one client id gives\nthe paragraph or the object below, where the entry in force and the\nentries it shadows are named.\n\nA client no pattern matches gets a row of `-` rather than no row. Silence\nwould make \"granted nothing\" and \"I never asked about that one\" the same\noutput, and those are the two answers somebody checking a fleet most needs\nto tell apart.\n\n**The header starts with `#`** so that it is skipped by the same rule that\nskips a comment in the list going in: the input is always the first column,\nso a kept run is asked again by cutting that column and piping it back, and\nthe header must not become one more question.\n\n**A configuration naming no `acl_file` is answered in one sentence**, in\nboth forms: it names the file it read and says that every authenticated\nclient may do anything. No rows at all. A row of `-` there would answer the\nmost permissive setting there is with the output this document defines to\nmean granted nothing, which is the reverse, and in the most emphatic form\nthe format has.\n\n#### `--json`, for a program rather than a person\n\n`--json` answers either form as JSON instead of columns. One object per\nline, compact, and no header - the keys are the header, and a `#` line is\nnot JSON:\n\n```sh\n$ printf 'device-7\\nnobody\\n' | saguin --acl examples/saguin.yaml --json\n{\"user\":\"device-7\",\"role\":\"device\",\"kind\":\"topic\",\"subject\":\"iot/+/health/device-7\",\"verbs\":[\"write\",\"read\"],\"denies\":[],\"publish_rate\":200,\"publish_bytes\":65536}\n{\"user\":\"nobody\",\"role\":null,\"kind\":null,\"subject\":null,\"verbs\":[],\"denies\":[],\"publish_rate\":null,\"publish_bytes\":null}\n```\n\n**A line per answer rather than one array**: an array is not valid until\nit is closed, so a run over a fleet would say nothing until it finished\nall of it, and `grep` stops working. `jq -s` collects an array for\nanybody who wanted one.\n\n**Where a column holds `-`, the field is `null`** - not the string `\"-\"`,\nwhich a consumer testing for absence would not recognise. `verbs` is a list\nand is `[]` rather than null, so that iterating it never has to test first,\nand so is `denies` - the feature denials a `broker: features` rule carries\n(*Taking a feature away*), and `[]` on every other rule.\n\n**`publish_rate` and `publish_bytes` are numbers, and `null` where\nnothing bounds one** - null is the permissive end there and the absent\nend in the four fields above it, there being no number that means\n\"unbounded\".\n\nOne client id gives one object instead, carrying what the paragraph form\nadds around the same grants - which entry applied, which patterns matched,\nwhich patterns the file holds, `grants_withheld`: the rules that grant this\npair nothing because the name they would put in holds `+`, `#` or `/`, as\nthe file writes them, and `client_id_allowed`: whether the entry's\n`client_ids:` admits the client id given, or `null` where none was. All\nfive are always present, because a shape that changes with the answer is\none a script has to branch on before it can read it; `pattern_applied` is\n`null` where nothing matched.\n\n**`patterns_matched` is not the answer to what a client gets**, and this is\nthe field to read instead. Every pattern that matches is listed, and one of\nthem is in force - a script testing whether a device matched its fleet\npattern will find that it did, while the entry actually supplying its roles\nis a narrower one beside it:\n\n```sh\n$ saguin --acl examples/saguin.yaml device-7 --json\n{\"user\":\"device-7\",\"acl_file\":\"/tmp/saguin-demo/acl.yaml\",\"pattern_applied\":\"device-7\",\"patterns_matched\":[\"device-*\",\"device-7\"],\"patterns_in_file\":[\"demo\",\"device-*\",\"device-7\"],\"grants\":[{\"role\":\"device\",\"kind\":\"topic\",\"subject\":\"iot/+/health/device-7\",\"verbs\":[\"write\",\"read\"],\"denies\":[]}],\"grants_withheld\":[],\"client_id_allowed\":null}\n```\n\nA configuration naming no `acl_file` answers `{\"acl_file\":null,\n\"everything_allowed\":true}`. An empty grant list there would be the JSON\nspelling of the row of `-` above, and would read as the reverse of the\ntruth.\n\n`--json` is how these two commands answer rather than a flag of its own:\ngiven without `--acl` or `--route` it is refused, rather than ignored while\na broker starts.\n\n### Where a topic lands: `--route`\n\n```sh\nsaguin --route /etc/saguin/saguin.yaml iot/depot/state/device-1\n```\n\n**Filters overlap deliberately, and which one holds a topic is settled by\na rule rather than read off the name.** So the broker answers it, against\nthe configuration, without publishing anything - for a topic, the channel\nthat holds it and the filter that claimed it; for a filter, every channel\na subscriber would be served from and what each would send.\n\nIt calls the functions the broker calls. A second implementation here\nwould agree on the day it was written and drift silently afterwards, which\nis the failure this command exists to prevent rather than to reproduce.\n\n**A list on standard input answers a line at a time**, in four\ntab-separated columns under a header naming them, the input first. Against\nthe configuration this repository ships, so that it can be run:\n\n```sh\n$ saguin --route examples/saguin.yaml \u003c fleet-topics.txt\n# topic-or-filter channel type why\niot/depot/events/order-1 events append iot/+/events/+\niot/depot/state/device-1 state latest iot/+/state/+\niot/depot/health/device-1 - broadcast -\niot/depot/# events append replayed\niot/depot/# presence latest current\niot/depot/# state latest current\niot/depot/# jobs queue excluded\niot/depot/# jobs__dlq append replayed\n```\n\n**`jobs__dlq` is in that list and is in no configuration file**, because a\nqueue derives a dead-letter channel. A wide filter reaching one is the sort\nof thing an operator finds out here rather than by reading, which is most of\nwhy the command exists.\n\nThe header starts with `#` so that it is skipped as a comment when a kept\nrun is asked again - cut the first column and pipe it back. A bare `#` is a\nlegal filter and is answered as one; a `#` with anything after it is a note.\n\nThe fourth column answers \"and so?\": for a topic, the filter that\nclaimed it - the line an operator goes and edits - and for a filter, what\na subscriber would actually be served. A filter reaching several channels\nis several lines rather than an invented summary.\n\n`awk -F'\\t' '$3 == \"broadcast\"'` is every topic no channel claims, and a\ndiff of two runs against two configurations is every topic a filter edit\nmoved. That is what the shape is for.\n\n**`--json` answers the same thing as JSON**, one object per line, on the\nsame terms `--acl` sets out above - a `-` column is a `null` field, there\nis no header, and one subject given as an operand gives one object instead,\ncarrying the routing table the paragraph form prints:\n\n```sh\n$ printf 'iot/depot/state/device-1\\niot/depot/health/device-1\\n' |\n saguin --route examples/saguin.yaml --json\n{\"subject\":\"iot/depot/state/device-1\",\"channel\":\"state\",\"type\":\"latest\",\"why\":\"iot/+/state/+\"}\n{\"subject\":\"iot/depot/health/device-1\",\"channel\":null,\"type\":\"broadcast\",\"why\":null}\n```\n\n**A blank line is skipped and a `#` line is a comment - except a line\nthat is only `#`**, which is the widest filter MQTT has and the thing\nsomebody most wants an answer about.\n\n### The operations listener\n\n`broker.operations` is the HTTP listener an operator reads: `/health`,\n`/metrics`, and the `/v1/operations` routes that answer *which one* where a\nmetric can only answer *how many*. RFC 0005 specifies what it serves and who\nreaches which route; this section is the keys that configure it.\n\n**`listen.tcp` and `listen.unix` are each the list-or-map shape \"Several\nlisteners of a kind\" gives every listen block**: the table below is a\nsingle door, named after its kind, and a second door of a kind is a list\nentry with a `name` of its own - RFC 0005 \"The operations listener\" has\nthe example. Every row below is already a per-door rule and stays one\nonce there are several: each door keeps its own `password_file`, TLS and\n`proxy_protocol`, and the loopback rule is asked of each TCP door in\nturn.\n\n| Key | |\n|---|---|\n| `listen.tcp.address` | `host:port`. Loopback only unless something there authenticates: a `password_file` naming who may read it - this listener's own, or the block's - or a client certificate required by `tls.client_ca_file` |\n| `listen.unix.path` | A socket file. Not reachable off the box, so it is the answer for a reader on this machine |\n| `listen.unix.mode` | Optional, `0660` if absent. The socket file's permissions, which are the access control |\n| `listen.tcp.tls` | Optional. A certificate and key, and `client_ca_file` for mutual TLS - the same block as an MQTT listener (*TLS on a listener*) |\n| `listen.tcp.password_file`, `listen.unix.password_file` | Optional. That door's own operators, instead of the block's `password_file` below |\n| `listen.unix.proxy_protocol` | Optional. The socket is fed by a proxy sending PROXY v2, which carries the Common Name it verified |\n| `min_scrape_interval` | Optional, `60s` if absent, and refused below `60s`. The shortest interval at which the metrics are recomputed |\n\n`listen` takes either entry or both, and at least one if the block is\npresent: a block that opens nothing is a configuration asking for\n`/metrics` and then not serving it anywhere. There is no default address,\nbecause a port opened by a default is a port nobody chose.\n\n**Omitting the block omits the listener.** No `/health`, no `/metrics`, no\nport. That is a deployment which is not scraped and does not want the\nsurface, and it is why `address` has no default to fall back on.\n\n**A TCP address that is not loopback has to authenticate somebody, and\nthe broker refuses to start otherwise**: a `password_file` - this\nlistener's own, or the block's - or `tls.client_ca_file` requiring a\nclient certificate (`require_certificate: false` does not count).\n`/metrics` names channels, volumes and consumer positions, and without\nthe refusal \"authentication later\" becomes for ever the first time\nsomebody edits the address to `0.0.0.0` because that made the scraper\nwork. Give it a password file, a tunnel, or the Unix socket.\n\n`password_file` is an absolute path to a file in Mosquitto's format,\nmanaged with `saguin --passwd`, holding the operators who may read\n`/metrics` - **not** the MQTT clients, which are a different set of people\nand a different file (RFC 0005). It is absolute for the reason every other\npath here is: which operators may read the metrics must not depend on where\nthe broker was started from.\n\n**A door may name its own instead**, and `listen.tcp.password_file` or\n`listen.unix.password_file` wins where it is written - the same rule an\nMQTT listener follows. The block's file is what a door that names none\ntakes, so the common deployment writes one file once and the deployment\nthis exists for writes two: a monitoring system on the port, whose\ncredential rotates with the estate's secrets, and a local agent on the\nsocket, which belongs to this machine. One file for both makes those two\nrotate together, which is what an operator splitting them is trying to\nstop.\n\n**A door that has no file anywhere has no credential**, which for the\nsocket is the ordinary arrangement - its file permissions are the gate -\nand for the port is what confines it to a loopback address. The rule below\nis asked of each door's own answer, so a listener naming its own operators\nmay bind a routable address, and a file on the *socket* says nothing about\nthe port.\n\n**There is no `allow_anonymous` beside it**, and one written under either\nlistener is refused at startup - the single place this differs from an\nMQTT listener. On the port `true` would open `/metrics` to whoever can\nreach it, and on the socket it would only strip the name off a reader\nalready getting in. A key whose settings are the default and a hole is\nnot a key.\n\n**`min_scrape_interval` is held in whole seconds**, so `500ms` is refused\nrather than rounded, and a scrape arriving sooner than the interval is\nanswered from the previous one - RFC 0005 \"The observer does not set the\ncost\" is why.\n\n**Sixty seconds is a floor as well as a default, and a smaller value is\nrefused rather than quietly raised.** A number the broker accepts and does\nnot honour is worse than one it turns down, because the scraper goes on\nasking every five seconds and the operator goes on believing it. Confluent\nCloud publishes its metrics on the same floor, and Prometheus's own default\nscrape interval is a minute - so this is the interval the ecosystem already\nassumes rather than a limitation of this broker. A deployment that wants\nfiner resolution than a minute wants a different instrument: the counters\nhere are cumulative, so a rate over a minute loses nothing but the shape of\na burst inside it.\n\n### Bridges\n\nA **bridge** is Sagüin as a client of another broker, in either direction:\nit subscribes at the far end and brings what arrives in, and it reads its\nown channels and publishes what it finds out. The peer is Mosquitto, EMQX,\nanother Sagüin, whatever is already there, and it needs no bridge\nconfiguration of its own - what arrives there is an ordinary MQTT client's\npublish, and what Sagüin reads is an ordinary subscription.\n\n**It is a translator rather than a copy.** The far end has never heard of\na channel or an offset, so everything that made a record Sagüin's is lost\non the way across and what arrives is a new record at an offset this\nbroker assigns. Disaster recovery is the database copy in RFC 0004, not\nanything a bridge does.\n\n**A bridged record enters this broker through its own publish path**:\nchannel resolution, size bounds, header limits and every reason code\nabove apply to it exactly as to a local client's, so a rule for a foreign\nfeed puts no restriction on channel type. **There is no second way into a\nchannel** - no bridge writes to storage behind the broker's back.\n\n**A record's identity crosses; its position does not.** A record leaving\ncarries its `saguin-id` as a User Property, and a Sagüin at the far end\nlifts that into the record it stores - so the same message has one identity\non both brokers (invariant 8), and a consumer can recognise a record it has\nalready seen. Nothing else of Sagüin's crosses: the offset is this broker's\nown count and the far end assigns its own, because what goes over the wire\nis MQTT and MQTT has no offsets.\n\nA record arriving from a **foreign** broker carries no `saguin-id`, because\nnothing there sets one - so it is a new message here and is given an\nidentity on arrival, as any publish is.\n\n**The RETAIN flag crosses, both ways** - and a bridge still ships records\nas they are written and never a stored pass, so a link restart re-ships no\nstored set. RFC 0003 \"Retained messages\" has the whole of it, the peer that\nadvertises no retained messages included.\n\n**One thing does cross twice, and MQTT is what requires it.** A QoS 1\npublish whose acknowledgement the upstream never saw is re-sent when the\nsession resumes - before any `SUBSCRIBE`, so no retained pass is involved -\nand a link that drops with an acknowledgement in flight produces exactly\nthat. Stored a second time it would take a second offset and carry no\n`saguin-id` either time, which is two records downstream that nothing can\ntell are one.\n\n**So the bridge remembers what it settled.** A record it stored and\nacknowledged is held as its packet identifier and a digest of its topic and\npayload, and an arriving publish that carries `DUP` and matches one is\nacknowledged again and not stored - the acknowledgement being the thing the\nupstream is missing. The memory is bounded by the `receive_maximum` this\nbridge advertises, which is exactly what limits what the peer may hold\nunacknowledged toward it, and the oldest is overwritten.\n\nIt is asked only of a packet carrying `DUP`, which is what makes a recycled\npacket identifier safe: an identical payload published again arrives as a\nfirst delivery, where MQTT forbids `DUP` [MQTT-3.3.1-1], so it is stored as\nthe new record it is.\n\n**It does not survive this process.** The memory is held in RAM, so a bridge\nthat restarts has forgotten what it acknowledged and a redelivery then\ncrosses as a second record. Between a link recovery and the next restart the\npromise holds; across a restart the bridge is at-least-once, as the rest of\nthis document already says delivery is.\n\n```yaml\nbridges:\n head-office:\n peer: tls://mqtt.example.com:8883\n client_id: vessel-07 # saguin's identity at the far end\n session_expiry: 1d # how long the peer holds the backlog\n receive_maximum: 20 # how much may be in flight to saguin\n ack_interval: 5ms # how long an acknowledgement may wait\n topics:\n - filter: fleet/+/telemetry/#\n topic: readings/telemetry/$1/$#\n direction: in\n - filter: alerts/#\n direction: both\n```\n\n**`filter` is what goes on the wire** as Sagüin's `SUBSCRIBE`, and it is an\nordinary MQTT topic filter - not a pattern evaluated locally over a broader\nsubscription. Asking the peer for more than is wanted and discarding\nthe rest pulls the difference across the link, which on the metered\nconnection a bridge exists for is exactly backwards.\n\n**`direction` is which way the rule carries, and it is required.** There is\nno default: an omitted key deciding whether this broker's records leave the\nbuilding is not something an operator could see in the file, so the rule is\nrefused and the refusal names the three words.\n\n| | What the rule does |\n|---|---|\n| `in` | subscribes to `filter` at the peer, and publishes what arrives into this broker |\n| `out` | reads what this broker holds under `filter`, and publishes it at the peer |\n| `both` | both of the above, on one filter |\n\n`out` reads a filter the way any client does: every `append` and `latest`\nchannel it matches, and every broadcast topic - **a queue is never crossed**,\nwhich is invariant 11 and the same answer a subscriber gets. An outbound rule\ndraining an `append` channel keeps a position in it, so a link that comes back\nresumes where it stopped rather than re-sending or skipping. **Each rule keeps\nits own**, named by its bridge, its filter and its topic: a peer refusing one\nrule's records holds that rule and no other, and retention passing it is\ncounted against that rule alone. A rule whose filter or topic is edited is a\nnew reader and starts where a new reader starts. A broadcast topic\nhas no position to keep, and what is published while the link is down is gone,\nwhich is what broadcast promises everywhere else. A `latest` channel is carried\nas its changes happen, never as a pass over its current state (RFC 0003): a\nchange made while the link is down waits in the rule, the newest per topic, and\ncrosses when the link returns.\n\n**An `out` rule publishes one record at a time per topic, and several\ntopics at once.** What it may have in flight altogether is\n`min(the peer's Receive Maximum, receive_maximum)` - the same key that\nbounds what the peer may have in flight *to* Sagüin, spent in both\ndirections.\n\n**Per topic is serial, and that is deliberate.** A record crossing a bridge\nis given a *local* offset where it lands, so at the far end currency is\narrival order: two writes to one topic overtaking each other would leave a\n`latest` channel there holding the older one for good, and a device reading\nit stale with nothing to correct it. Order across *different* topics is not\npromised and never was - RFC 0003 promises order per topic.\n\nSo **a stream on a single topic still costs one round trip per record.**\nThat is the residual rather than a shortfall: it is what per-topic order\ncosts, and no window can buy it back.\n\n**What it costs, measured rather than reasoned.** Two Sagüins with a\nlatency-injecting proxy between them, one `out` rule, records already\nwaiting in the channel:\n\n| Round trip | One topic | Eight topics |\n|---|---|---|\n| ~20µs (loopback) | 10,971/s | 11,721/s |\n| ~5ms | 148/s | 552/s |\n| ~10ms | 89/s | 299/s |\n\nOn one topic the rate is one record per round trip and nothing else: a 20ms\nlink - an ordinary satellite or cellular figure, and the deployment a bridge\nexists for - carries about 47 records a second however fast either broker\nis. Spread over eight topics the same link carries three to four times\nthat, because the window is spent across them.\n\n**Loopback hides all of this**, which is worth stating for anyone measuring:\nat 20µs a round trip costs nothing and both columns look like no limit at\nall. The shape only appears once the link has latency, which is where\nbridges live.\n\nSagüin keeps in-order delivery per topic and spends the window across\ntopics.\n\n**What a window costs, and what bounds it.** Duplication across the hop\nalready existed: a link that drops with records unacknowledged re-sends them\non the next connection, because the position never moved past them. A window\nwidens that bound from one record to the window, and only on a refusal - a\npass cannot know which record the peer will refuse until the ones behind it\nare already in flight, and those are offered again on the next pass. The\nfirst refusal stops the pass issuing anything further, so what may be sent\ntwice is what was in flight and not the whole batch.\n\nThe bound is therefore `receive_maximum`, which an operator sets. And the\nrecord's identity crosses with it - `saguin-id`, which a receiving Sagüin\nlifts into the record it stores - so a consumer that deduplicates on message\nidentity sees each record once whatever the link did. The position itself is\nunaffected: it advances only through the contiguous run of acknowledged\noffsets and stops at the first the peer did not take, so nothing is ever\nskipped.\n\n**`topic:` is required on an `in` or an `out` rule, and refused on a\n`both`.** The two halves of that are one rule read from either side. A\none-way rule states the whole topic its records take at the other end, so\nleaving it out would be a rule that does not say where anything lands. A\n`both` rule cannot carry one at all: a template is a one-way rewrite with\nnothing to invert it, so `both` plus a template would mean \"rewrite going\nout, do not coming back\" - two rules wearing one name. `both` therefore maps\na topic to itself, and that identity is what makes it reversible.\n\n**`topic` is the whole topic a record takes, and nothing else decides where\nit lands.** A channel is a topic filter rather than a prefix, so there is no\nsuffix that generally sits inside one - `iot/+/+/events/#` has nothing to\nhang a suffix off. The topic resolves exactly as it does for a publish from\nany client: to whichever channel's filter matches it, or to broadcast when\nnone does.\n\n**There is no key naming the channel, and that is deliberate.** A client\nnever has to know where the channels are - it publishes a topic - and a\nbridge is a client. A rule that named one would be a second routing table\nbeside the channel filters, and two routing tables disagree eventually.\n`--check-config` prints what each rule reaches, which is where an\noperator checks where a rule's records will land:\n\n```\nsaguin.yaml: ok\n alerts latest\n readings append\n bridge \"head-office\", in \"fleet/+/telemetry/#\": lands in readings (append)\n bridge \"head-office\", both \"alerts/#\": lands in alerts (latest)\n```\n\nA filter that crosses a queue says so too, because nothing else ever will -\nsuch a rule is served everything it matches *except* the queue, exactly as\nany subscriber's is (invariant 11), so there is no refusal at startup and no\nline at runtime:\n\n```\n bridge \"head-office\", both \"iot/#\": lands in events (append), jobs__dlq (append), state (latest); crosses jobs (queue), never served\n```\n\n**A rule may not publish into the reserved ` RFC 0002 - Channels and configuration | Sagüin documentation space**, where MQTT keeps its\nown topics and Sagüin keeps a queue acknowledgement and a consumer's seek. A\nbridge able to publish there would forge them on behalf of whatever it\ncarries. A template cannot express it - a ` RFC 0002 - Channels and configuration | Sagüin documentation always begins a substitution,\nso `$saguin/…` is refused as a ` RFC 0002 - Channels and configuration | Sagüin documentation that is neither `$#` nor a wildcard\nnumber - so what is actually stopped is the topic a substitution builds out\nof what the peer published, and it is dropped with a warning. An inbound\n`filter` beginning `$saguin/` is refused at startup for the same reason.\n\n**`topic` is that whole topic, with substitutions**: `$1`, `$2` … for the\nfilter's `+` levels in order, and `$#` for its `#` tail.\n\n| Filter | Topic | `fleet/vessel-07/telemetry/hold/psi` becomes |\n|---|---|---|\n| `fleet/+/telemetry/#` | `readings/telemetry/$1/$#` | `readings/telemetry/vessel-07/hold/psi` |\n| `fleet/+/telemetry/#` | `readings/telemetry/$#` | `readings/telemetry/hold/psi` |\n| `fleet/#` | `readings/$#` | `readings/vessel-07/telemetry/hold/psi` |\n\nThe numbered captures are what a prefix strip cannot do: they **reorder**,\nso `fleet/\u003cvessel>/telemetry/\u003cmetric>` can become\n`telemetry/\u003cvessel>/\u003cmetric>` rather than merely a shorter version of\nitself.\n\nThere are no regular expressions, and the reason is worth stating because\nthe question will be asked again. A regex cannot be sent to a peer\nbroker, so a regex filter forces a broad subscription and local matching,\nwhich is the waste above. It also admits patterns that look right and are\nnot: `fleet/(.)+/telemetry/(.)+` compiles, runs, and captures **one\ncharacter** - `$1` is `7` and `$2` is `p` - so every record lands on a\ntopic made of stray characters and nothing reports it. A `+` is one topic\nlevel by construction and cannot mean anything else.\n\n**A `#` matches zero levels as well as many** (MQTT 5 section 4.7.1.2), so\n`fleet/#` reaches the bare `fleet`. The tail is then empty and the\nseparator that was joining it goes with it, so `telemetry/$1/$#` gives\n`readings/telemetry/vessel-07` rather than a topic with an empty last\nlevel. Where the whole suffix would be empty there is nothing to publish -\na channel topic has a non-empty suffix - and the record is dropped with a\nwarning rather than published somewhere unexpected.\n\n**Rules are tried in the order they are written and the first match wins.**\nA message matching no rule is dropped with a warning. That makes the order\nof the list significant, which is worth knowing before a merge reorders two\nentries and changes what the broker does without changing what any line\nsays.\n\nMatching is a linear scan, and the cost is the reason it is allowed to be.\n`BenchmarkInboundMatch` and `BenchmarkInboundMatchMiss`, at `-benchtime 2s\n-count 2`:\n\n| one rule, against one message | Ryzen 7 260 | MacBook Pro M1 Pro |\n|---|---|---|\n| that does not match | 54ns | 64ns |\n| that matches | 143ns | 162ns |\n\nOnly a message that matches *nothing* runs every rule, so fifty rules is\nunder 3us on traffic nobody asked for. An index earns its keep a long way\nabove that.\n\n#### Loops, and what stops one\n\n**A record that arrived over one of this broker's own bridges is never\nsent out over another** - a blanket mark on the record, not a comparison\nagainst the bridge's own name (RFC 0003 \"Retained messages\" has why, and\nre-origination for a broker meant to relay).\n\n**Per link, that is the whole of it.** Two brokers, one `direction: both`\nrule, and nothing is duplicated: No Local [MQTT-3.8.3-3] stops the peer\nechoing this bridge's own publish back down the connection it arrived on, and\nthe mark stops what the bridge pulled in from going back out. Against a peer\nthat ignores No Local the worst case is one duplicate per record - bounded,\nand never a loop.\n\n**Across links it is topology rather than mechanism, and the obligation is\nthe operator's: the bridge graph must be a tree.** The mark cannot cross the\nwire. A record a peer *pushed* here arrives as an ordinary client publish -\nwhich is the whole of the promise that the far end needs no bridge\nconfiguration - and nothing local can tell it from a record a device\npublished. So relay depends on which way each link was dialled:\n\n| The record was | Onward over another bridge |\n|---|---|\n| pulled in by a rule here | stopped |\n| pushed here by a peer's rule | forwarded |\n\nEdge to hub to cloud therefore works when the links dial toward the cloud,\nand a hub that pulls from the edge does not relay onward - which is the case\nfor re-origination: a broker meant to relay publishes the record again as its\nown, the way dead-lettering already does.\n\nA cycle loops. Three brokers each dialling the next multiplied one record\nforty thousand times in three seconds, and no broker in the ring can decide\nlocally that it should not have. Mosquitto states the same obligation for the\nsame reason. Sagüin cannot refuse a topology it cannot see: it knows its own\nrules and nothing about the peer's.\n\n#### What running a bridge means\n\n**Sagüin does not acknowledge the peer until its own store has taken\nthe record.** That is the rule the rest of this follows from. A channel at\nits `max_bytes` answers `0x97`, and a bridge that had already sent its\nPUBACK would have discarded a record the peer believes it delivered -\nacknowledged and lost, which this document refuses everywhere else, one hop\nfurther out than Sagüin can otherwise produce it.\n\n**So a full local channel becomes backpressure on the peer\nsubscription, and that is intended.** Sagüin subscribes with a Receive\nMaximum, so a peer that stops being acknowledged stops sending that\nmany records later. It is worth saying out loud because of what it means\nfor the other broker: an inbound bridge into a channel nobody is draining\nwill stop draining a broker other people are using. Against mosquitto\n2.1.2, a subscriber asking for a Receive Maximum of 5 and never\nacknowledging receives exactly 5 of 50 published.\n\n**The backpressure holds the peer's records, and nothing else.** The bridge\nstores what arrives on a worker of its own, in order, rather than on the\nconnection that brought it, so a channel refusing a record never keeps the\nbridge from noticing that its link has gone - `saguin_bridge_connected`\nfalls when the link does, and it reconnects. A QoS 1 or 2 record waits\nunacknowledged, as above. A QoS 0 record has no acknowledgement to hold\nback, so while the channel refuses, those past one Receive Maximum of\nwaiting records are dropped and counted\n(`saguin_bridge_unstored_total{cause=\"queue_full\"}`, RFC 0005):\nat-most-once, which is what their publisher asked for, and mosquitto's\nrule for QoS 0 over a full queue.\n\n**A record that could never be accepted is dropped rather than retried.**\nA refused record is published again until it gets in, without a limit,\nbecause a full channel and a storage that is not working can both stop\nbeing true. A record that carries more topic or more headers than Sagüin\nallows cannot, and retrying it would hold the link for ever behind\nsomething that will never move. So it is dropped with a warning, and it is\nthe fourth of them. Each is counted under its cause in\n`saguin_bridge_unstored_total` (RFC 0005), because each was finished with\nat the peer and is lost at this hop:\n\n| Dropped, with a warning | Because | `cause` |\n|---|---|---|\n| No rule covers the topic | Nothing asked for it | `no_rule` |\n| A rule covers it and produces no topic | A `#` matched zero levels, and a channel topic has a non-empty suffix; or the peer sent an empty topic, which nothing may publish | `unmappable` |\n| A rule produces a topic in the reserved ` RFC 0002 - Channels and configuration | Sagüin documentation space | A bridge may not forge Sagüin's own control topics on behalf of what it carries | `unmappable` |\n| Sagüin would refuse the record whatever the channel | It can never be accepted, and holding it stops everything behind it | `never_accepted` |\n\n**An outbound rule drops for the middle two of those, and says so the\nsame way** - the check is one piece of code for both directions, so a\nrule that would publish `$saguin/…` at the peer is refused here whatever\nthe peer is. The record stays at its offset under this channel's\nretention, and the drop is counted and logged: nothing at the far end can\nnotice a record that never arrived.\n\n**An inbound bridge is only as durable as the broker it reads from.**\nSagüin connects with a persistent session, so the peer holds messages\nwhile the link is down - the buffer is somebody else's broker, bounded by\ntheir rules. Mosquitto's `max_queued_messages` defaults to 1000: a\nsession offline while 1200 are published comes back with exactly 1000,\nthe peer logs it, and Sagüin cannot see it. **Sagüin's guarantees do not\nextend one hop out**, and somebody will assume they do. An *outbound*\nrule keeps the property: what it drains is this broker's own channel,\nbounded by retention rather than by somebody else's session queue.\n\n**A duplicate on the link is recognisable between two Sagüins and not\notherwise.** MQTT is at-least-once, so a lost acknowledgement makes the\npeer re-send; between Sagüins the copy carries the same `saguin-id` at\nits own offset, so a consumer keeping identities can tell (invariant 8).\nA foreign broker's record carries none, so Sagüin mints one and the\nduplicate is a separate message nothing can match: on an append channel\none a consumer cannot tell from a record, on a **queue** two jobs, done\ntwice, outside protections that are about one record and two workers.\n\n**A rule may target a queue and this is the cost of it**, stated here as a\nwarning rather than as a preference. A queue is the one channel where the\nmitigation below does not exist: pointing the rule at a `latest` channel\nworks because the same record replaces the value already there, and work\ndoes not collapse - two jobs are two jobs. What is left is `max_attempts`\nand the dead-letter channel, which bound how often a duplicate is *retried*\nand catch what fails, and neither of them can tell a duplicate from a job.\nSo a fleet whose work arrives over a bridge needs idempotent workers, which\nis what at-least-once already asked of it, and needs them for a reason that\nis the link rather than the queue.\n\n**The largest duplicate is not a lost acknowledgement, it is a lost\nsession.** A lost acknowledgement re-sends one record, occasionally. A\nresubscribe re-sends the peer's **entire retained set**, because that is\nwhat a broker answers a subscription with - so every reconnect that finds\nSession Present = 0 copies all of it in again, deterministically. That is\nthe case the bridge already warns about, and it is what an edge box down for\nlonger than `session_expiry` produces. On an `append` channel those\naccumulate for the life of the deployment.\n\n**Pointing the rule at a `latest` channel is the mitigation**, and it is the\none an operator can choose in advance: the same records replace the values\nalready there, so the channel holds what the peer holds however many\ntimes the link is rebuilt. It is the right shape for the traffic anyway,\nsince a retained message *is* a current value per topic. An `append` channel\nis for a stream of events, where re-reading the retained set is a duplicate\nwith nothing to collapse it.\n\n**One log line when the link goes and one when it returns**, saying how\nlong - not one per attempt, which scrolls a bad connection over the record\nof everything else that happened. Reconnection is capped and jittered.\n\n#### The three keys that tune a link\n\nThey are per bridge, because the thing they describe is the far end and no\ntwo peers are alike: a bridge to a broker on the same LAN and one to a\nbroker over a metered satellite link want different answers to all three.\n\nEach describes the subscription: how long the peer holds Sagüin's\nsession, how much it may have in flight *to* Sagüin, and how long Sagüin\nmay wait before acknowledging it.\n\n| | Default | What it decides |\n|---|---|---|\n| `session_expiry` | `1d` | How long the peer keeps Sagüin's session, and so how long an outage it holds a backlog across |\n| `receive_maximum` | `20` | How many records the peer may have in flight to Sagüin at once |\n| `ack_interval` | `5ms` | How long a record Sagüin has already stored may wait before the peer is told |\n\n**`session_expiry` spends somebody else's storage** - the peer holds the\nbacklog, so raising this asks another operator to hold a week of records\nfor a broker they may not know exists, and it only stops the *session*\nbeing discarded before the queue is. Lower it, and a Sagüin down longer\ncomes back to Session Present = 0, warned about, having lost whatever was\nqueued. A day is the outage a bridge exists to survive.\n\n**`receive_maximum` is the backpressure, and it bounds what a bridge holds\nunacknowledged** (invariant 13). Sagüin acknowledges a record only once its\nown store has taken it, so a channel that stops accepting stops the\nacknowledgements, and the peer stops sending this many records later.\nSetting it to 1 makes the link a round trip per record. Setting it high\nmakes a full channel take that much longer to become backpressure, and holds\nthat many records in memory in the meantime.\n\n**`ack_interval` is the one that is not obviously a knob until the\narithmetic is written down.** An acknowledgement is not sent the instant\nSagüin stores a record; it is marked, and a ticker sends what has been\nmarked. Since the peer will not send past `receive_maximum`\nunacknowledged records, **the two together are a ceiling of\n`receive_maximum / ack_interval`**.\n\nIt is a ceiling and not a rate, and the difference matters at the short\nend. Measured at `receive_maximum: 20` against a mosquitto peer on the\nsame machine: 200 records published to it at QoS 1 from one connection,\ncounted as they reach a subscriber of the memory-backed channel they land\nin - so the publisher is part of what is measured, as it is on any link:\n\n| `ack_interval` | the ceiling says | Ryzen 7 260, mosquitto 2.0.22 | M1 Pro, mosquitto 2.1.2 |\n|---|---|---|---|\n| `50ms` | 400/s | 412–413/s | 444–454/s |\n| `5ms` | 4,000/s | 2,800–3,000/s | 3,600–4,300/s |\n| `1ms` | 20,000/s | 10,300–11,800/s | 5,000–5,700/s |\n\nAt `50ms` the interval is the whole story. Below it something else is the\nslower half, and shortening the interval further buys less each time. So\nthe arithmetic says where the ceiling is, not what the link will do, and\nthe useful reading is that the default is comfortably off its own ceiling\nwhile a much longer interval would not be.\n\nIt is also the window in which a record Sagüin has stored is not yet\nacknowledged at the peer. A link cut inside that window makes it\nredeliver, and that is the duplicate above: a second record at its own\noffset, carrying the original's `saguin-id` where the peer is a Sagüin and\na fresh one where it is not. Lowering it narrows that window and raises\nthe ceiling, and costs a timer wake-up per interval for the life of the\nbridge, on hardware where that is not free.\n\nIt is one of the five keys in this file that may be given a value below a\nsecond - a `sqlite` provider's `publish_commit_interval` and\n`flush_interval`, `broker.session.ack_commit_interval` and\n`limits.write_timeout` are the others - and they are why the duration form\nadmits `ms` at all.\n\n**What a bridge does not do yet.** It authenticates with a certificate and\nnothing else: `cert_file` and `key_file` are the certificate Sagüin presents\nwhen the peer asks for one, and there is no username and no password. A\nbridge cannot be given either, and the password file Sagüin reads is not the\nplace to look for one - that file is for connections arriving at this\nbroker, and a bridge is Sagüin connecting outward to somebody else's.\n\n### Where a reader starts: `start`\n\nWhere a subscriber with no stored position begins reading an `append`\nchannel. Two values:\n\n| Value | A reader with no position is served |\n|---|---|\n| `floor` | the default - everything the channel still holds, then what follows |\n| `tail` | only what arrives after it subscribed |\n\n**It is an `append` key and is refused on the other two types**, rather\nthan accepted and ignored. A `latest` subscriber is always served current\nstate, which is what that channel type *is*; a `queue` has no per-consumer\nposition at all, because work is claimed rather than read from a place. A\nkey that quietly does nothing on two of the three is a key somebody sets\nand then trusts.\n\n**What decides it is what the channel holds** - a channel of events wants\nthe floor, a channel of commands wants the tail - and the operator knows\nwhich when the broker cannot. One answer for every reader whatever\nprotocol it speaks, and the rest, history on a `tail` channel by a seek\nincluded, is RFC 0003 \"Where a subscription starts\".\n\n```yaml\nchannels:\n commands:\n type: append\n filter: iot/+/+/commands\n start: tail # a rebooting fleet must not act on old orders\n```\n\n### Bounds on what a channel holds\n\nThree keys, and the two size bounds are not two spellings of one thing:\n\n| | `retention_period` | `retention_bytes` | `max_bytes` |\n|---|---|---|---|\n| `append`, and a queue's `__dlq` | yes | yes | yes |\n| `latest` | yes | - | - |\n| `queue` | - | - | yes |\n| `broker.retained` | yes | - | - |\n| a provider | - | - | yes |\n\n`retention_bytes` **removes** the oldest records to stay under. `max_bytes`\n**refuses** the publish that would exceed, with `0x97`. The difference is\nwhat happens to a producer: retention deletes behind it and never says so,\nwhich is right for telemetry that ages out and wrong for anything an\nauditor will ask about, while `max_bytes` says no and keeps everything,\nwhich is right when losing a record is worse than refusing one. The one\nthing a provider gives up to make room is the broadcast log of the\nsessions it holds, never a channel's record (\"Every session's state:\n`broker.session`\").\n\n**What a size bound counts is a record's size, not the file's.** A\nchannel's `retention_bytes` and `max_bytes`, and a memory provider's\n`max_bytes`, count each record as its payload, its topic, its message id,\nand the key and value of each of its User Properties (`store.RecordSize`),\nand a memory provider counts the sessions it holds as well, each as the\nmemory the broker holds for it (\"Every session's state: `broker.session`\").\nA sqlite provider's `max_bytes` is different: it bounds the file itself, in\npages (RFC 0004 \"A sqlite provider's bound is SQLite's own\"). So a sqlite\nfile is larger than what its channels count. Measured on one append channel,\nafter a clean close, a record took about 75 bytes more on disk than it was\ncounted: 280 bytes against 205 at a 128-byte payload (1.37x), and 16,963\nagainst 16,461 at 16KB (1.03x). While the broker runs the write-ahead log\nadds to that (RFC 0004 \"WAL, and `synchronous=NORMAL`\"), and pages that\nretention frees are reused rather than returned, so the file does not shrink\n(RFC 0004 \"Reading it while the broker runs\"). Size a sqlite provider's\n`max_bytes` above the sum of its channels' bounds by that margin.\n\n**A `latest` channel takes neither size key**, and the reason is its shape\nrather than its cost. It holds one value per topic, so it is not a history\nthat grows - what grows is the number of topics, and the tool for a topic\nthat has gone quiet is the period. A byte cap there would answer a question\nthe channel does not pose.\n\n**A queue takes `max_bytes` and never retention.** An age- or size-based\nrule that deletes unacknowledged work is eviction of unresolved work,\nwhich is never permitted (invariant 2). A queue's size bound is therefore\nbackpressure and not a limit: resolution is what frees the space, so a\nqueue at its bound refuses new work with `0x97` until its workers catch\nup, and recovers on its own.\n\nA queue's dead-letter channel is an ordinary `append` channel and takes\nordinary retention - but it cannot be configured directly, because derived\nnames are part of the topic contract. Its two keys live in the queue's\nblock instead, spelled the same with a prefix: `dlq_retention_period` and\n`dlq_retention_bytes`. They are copied onto the derived channel, which\nalready inherits the queue's storage provider the same way.\n\n### Retention is stated once, and may be `none`\n\n`broker.storage.default_retention_period` and `default_retention_bytes`\nare what a channel gets when it does not say. Both are **required**,\nexactly as `broker.storage.default` is, and both accept `none`.\n\nThere is no built-in figure, because both possible defaults are wrong in\nthe way this file already refuses elsewhere. A built-in period makes a\ndurable, replayable log quietly mean \"the last three days of one\", and a\nreader who did not exist yesterday finds yesterday's events missing with\nnothing to tell them - the retention floor only reports the gap to a\nconsumer that already had a position. A built-in `none` makes every\nchannel grow until the disk does not, which is invariant 13's failure with\nno defined behaviour at the bound. So the operator writes the number, once\nfor the broker, and a channel that differs overrides it.\n\n**The default applies to `latest` channels too, and means something else\nthere.** On an `append` channel a period drops old events. On a `latest`\nchannel it deletes the *current value* of any topic that has gone quiet\nfor longer than the period, so a device reporting weekly loses its state\nunder a three-day default. That is deliberate - a stale reading is worse\nthan none, and RFC 0003 says so where `latest` expiry is described - but\nit is the one place where one rule across two channel types produces two\ndifferent outcomes, and a fleet with slow devices will meet it. Such a\nchannel states its own period.\n\nSet the period to what makes a value worthless rather than to what makes\nthe disk comfortable; the consumer's half - nothing older than the period\nis served, and every value says when it was set - is RFC 0003's.\n\n### Retained messages on a broadcast topic\n\nA client may ask for a message to be kept as the current value of its\ntopic - the RETAIN flag. What each channel type answers it is under\n*Publishing*; on a broadcast topic Sagüin keeps the value in the retained\nstore, which every broker has, and one optional block changes where it is\nkept and for how long:\n\n```yaml\nbroker:\n retained:\n storage: local # defaults to broker.storage.default\n retention_period: none # the default\n```\n\n**Absent means the defaults, never refused.** The store is bounded by its\nprovider's `max_bytes`, like everything else on that provider, so a client\nsending the flag cannot grow it past what the operator wrote. Taking\nretained values away from a client is the `acl_file`'s job (\"Taking a\nfeature away: `broker: features`\"); there is no broker-wide switch.\n\n**It is one store, and it is not a channel.** It claims no prefix, has no\nname, and holds only topics that no channel claims - one store rather\nthan channels a client's publishes could bring into existence unbounded,\nwhich the configuration file would not describe.\n\n**What it holds, and what it hands back.** One current value per topic,\nwith a zero-length payload deleting the entry, which is MQTT's own\nconvention. A subscriber is sent the stored value of every topic its\nfilters reach at the moment it subscribes, carrying the RETAIN flag so that\nit can tell state it is catching up on from an update that has just\nhappened. A retained publish is *also* delivered live to whoever is\nsubscribed at the time, because it is still a broadcast.\n\nSagüin delivers those values itself, from this store, and the server's own\nretained store is left holding nothing: a broadcast publish passes through\nit and is taken straight back out, and a channel's record never enters it\nat all. That is the same arrangement `latest` channels already use, and for\nthe same reason - a subscriber is owed what is current when it subscribes,\nwhich is not what a fan-out at publish time delivers.\n\nThe value goes back out as it was published. It carries no `saguin-id` and\nno `saguin-offset`: a broadcast message has no record identity and no\nposition, and stamping one on the way out would make broadcast look like a\nchannel to whichever client happened to subscribe late.\n\n**`retention_period` defaults to `none`**: a period deletes the current\nvalue of any topic quiet for longer than it, so a device reporting weekly\nwould lose its state under any defaulted figure. Expiry that belongs to\none value is its publisher's Message Expiry Interval, which applies to\nthe stored value; an operator who wants a blanket period states one.\n\n**The two clocks delete independently, and the value dies at whichever\nruns out first.** A publisher's interval removes the value whether or not\na period is configured, which is what makes the default above safe; a\nperiod removes a value whose publisher set no interval, which is most of\nthem. What a subscriber is told is the publisher's countdown alone,\ndecremented by the wait - the operator's period is not a client's\nbusiness, and a value already past its expiry is not served at all.\nRFC 0003 \"Retained messages\" has the delivery half.\n\n**Size is the provider's `max_bytes`, and the answer at the bound is a\nrefusal**, `0x97 (Quota exceeded)`, with what is already stored untouched:\ndeleting one device's current state to make room for another's would be a\nbroker losing something still true, silently. The store takes neither\n`retention_bytes` nor a `max_bytes` of its own, for the reason a `latest`\nchannel takes neither.\n\n**The `CONNACK` says Retain Available 1**, to every client. A client the\n`acl_file` denies `retained` is told 1 as well, because channels still take\nthe flag from it, and MQTT gives one answer for the whole broker with no way\nto say \"on these prefixes\".\n\n**A topic that becomes a channel's is moved or dropped at startup, and is\nnever held in both places.** Configuration changes only at startup, so that\nis where it is settled, for every stored topic a channel now claims:\n\n- Into a **`latest`** channel the value is **moved**, and the startup line\n says how many went where. The two are the same thing under different\n names - one current value per topic - so nothing is lost and nothing\n changes meaning.\n- For an **`append`** channel, a **`queue`**, or a dead-letter channel the\n values are **discarded**, and the startup line names the channel and\n counts them. Keeping them would not be a move: in an append channel a\n current value becomes an event that never happened, carrying a timestamp\n older than records already in the log, which is the ordering every\n consumer and every seek by time depends on; in a queue it becomes work\n nobody sent.\n- A move the destination cannot take - its provider is at `max_bytes` - is\n a **startup error** naming the channel, the provider and how many values\n it was, rather than a broker that starts having quietly dropped state an\n operator asked it to keep.\n\n### Exactly-once publishes: `broker.qos2`\n\nQoS 2 transfers ownership of the message to the broker at the `PUBREC`,\na full round trip before the client is told it is done, so a broker\noffering it holds messages nobody has finished sending. Every Sagüin\noffers it, and one optional block changes how many one client may hold and\nfor how long:\n\n```yaml\nbroker:\n qos2:\n max_inflight_per_client: 20 # the default\n expires_after: 5m # the default\n```\n\n**Absent means the defaults.** There is no `qos2: false` and no broker-wide\nswitch: taking QoS 2 away from a client is the `acl_file`'s job (\"Taking a\nfeature away: `broker: features`\").\n\n**Each held message sits in the store of the channel it is for**, or the\nbroadcast log's for a broadcast, inside that store's provider's `max_bytes`\nand inside `max_inflight_per_client`, which counts one client's held\nmessages across every channel and broadcast - never in memory nothing\nbounds. The release is one operation in the store that keeps the record,\nso a refusal or a crash cannot fall between taking the message out of one\nstore and writing it to another.\n\n**The record is written when the `PUBREL` arrives**, and RFC 0003 has\nwhat that buys. A held message outlives a restart as far as the provider\nof its channel, or of the broadcast log, does and is handed back to the\nsession that sent it; an exchange whose session did not come back is\ndropped at that start and counted.\n\n| Key | Default | |\n|---|---|---|\n| `max_inflight_per_client` | `20` | How many exactly-once publishes one client may have unreleased at once. Past it the next is refused on its `PUBREC` with `0x97 (Quota exceeded)`, before ownership is taken, and the connection is left alone. **Per client**, so one publisher that opens exchanges and never releases them stalls itself rather than taking room from publishers that are behaving. A count rather than a size, because `limits.max_message_size` already caps what one can be |\n| `expires_after` | `5m` | Drops a publish that has waited this long for its `PUBREL`, and the release arriving afterwards is answered `0x92`. A connected client releases within a round trip, so what this covers is one whose link dropped after the `PUBREC` and comes back to the same session to finish. Without it, such a message waits until the session ends, which `limits.max_session_expiry` puts at 30 days |\n\nThere is no `max_bytes` here, for the reason `broker.retained` has none: a\nheld message counts against the **provider's** `max_bytes` - its\nchannel's, or the broadcast log's - and what bounds one client is\n`max_inflight_per_client`. Writing the key is a startup error rather than\na bound somebody believes they set.\n\n`saguin_qos2_held`, `saguin_qos2_abandoned_total` and\n`saguin_qos2_max_inflight_per_client` are what an operator watches (RFC\n0005), and the second of those is the only place a publisher that opens\nexchanges and never finishes them is visible at all.\n\n### What a shared group is owed: `broker.share`\n\nA shared group whose members are all away is owed the deliveries that\nmatched it while nobody was there to take them. They stay in the broadcast\nlog, in the provider `broker.session.storage` names, behind a cursor the\ngroup keeps with the sessions (RFC 0003 \"Broadcast\"), so a backlog keeps\nacross a restart what the sessions keep. One optional block says how long\none waits:\n\n```yaml\nbroker:\n share:\n expires_after: 1h # omitted, a backlog lives as long as a member session\n```\n\n| Key | Default | |\n|---|---|---|\n| `expires_after` | none | Drops a backlogged delivery that has waited this long. Omitted, a backlog lives exactly as long as a member session that could collect it, which is the conservative default: the alternative is a broker quietly discarding work nobody asked it to bound |\n\n**A group holds nothing unless one of its members has a session that\noutlives its connection** - a backlog for an all-clean-start group is\nmemory nothing will ever collect, so those deliveries are dropped and\ncounted. **Only QoS 1 and 2 are held**: a QoS 0 delivery is owed to\nnobody by anything.\n\nWhat one group's backlog may cost is `limits.session_queue_bytes`, and when\nit is reached the oldest goes, as a session's does - a group is owed\nmessages the way a session is, and bounding it by a second figure that had\nto agree with the first is two numbers where one will do.\n\n### Every session's state: `broker.session`\n\nA session is what MQTT keeps for a client between packets: its\nsubscriptions, the messages it is owed or has in flight, and the Will it\narmed. One optional block says which provider keeps it, for every session,\npersistent or not:\n\n```yaml\nbroker:\n session:\n storage: local # defaults to broker.storage.default\n ack_commit_interval: 200ms # the default\n```\n\n**Each kind of data has one storage key, named after the data**, and this is\nthe one for sessions: a channel's records go where the channel's `storage`\nsays, retained messages where `broker.retained.storage` says, and session\nstate here, whatever the client asked for.\n\n**What survives is what the provider survives**: a persistent session on a\n`sqlite` provider survives a crash, on a memory provider a graceful\nshutdown. What comes back with a session and what ends with it is RFC 0003\n\"Sessions\".\n\n**What it costs is one write per publish.** A QoS 1 or 2 broadcast owed to\nany session that outlives its connection is written to the provider's\nbroadcast log once, however many sessions are owed it, and each session\nkeeps only its cursor and its window (RFC 0003 \"Broadcast\"). Crash survival\nor throughput is the operator's pick of provider; QoS 0 has nothing to keep\nand costs nothing either way.\n\n**`ack_commit_interval` is how long a connected session's acknowledgements\nmay wait to be taken for writing.** A channel consumer's position and a\nbroadcast session's acknowledgements are both taken for writing at most\nthis long after they move, however much the session is still owed, off the\nconnection's read loop, so no client's acknowledgements wait on the write;\nthey are stored when that write reaches the provider. On a `sqlite`\nprovider the write waits behind those already queued for its one\nconnection - the writes a client's packet is waiting on first, the rest in\nturns (RFC 0004 \"Group commit\") - so under a storage backlog, a reconnect\nstorm or a slow disk, they are stored later than the interval; on a\n`memory` provider there is no queue. It is what an unclean stop sends\nagain: invariant 18's first exception, never more than a session's Receive\nMaximum messages of the broadcast log, and RFC 0003 \"Broadcast\" and\n\"Durable consumers\" say what that is on the wire. A clean `DISCONNECT` and\na graceful stop store them whatever this says.\n\n| | `ack_commit_interval` |\n|---|---|\n| **Absent** | `200ms` |\n| **Accepted** | a duration from `10ms` to `1s` |\n| **Refused** | `0` and anything negative; `none`, since nothing stops acknowledgements being stored; below `10ms`; above `1s`; anything unparseable |\n\nPast a second, the last moment an unclean stop sends again is seconds of\nduplicate traffic for every busy session, and acknowledgements wait that\nlong in memory and in the store. Below ten milliseconds nothing is gained:\nstoring after every batch measured no faster than the default.\n\n**A Will is session state, and it is kept here from the `CONNECT` that\narmed it** (section 3.1.2.5). It has to survive a restart because of the\nWill Delay Interval: a Will waiting out its delay is an announcement\nnobody has made yet, and a broker restarting inside that wait would\nrestart into silence - the device gone, and nothing ever said.\n\n**A Will the provider has no room for refuses the connection**, `0x97 (Quota\nexceeded)`, with a Reason String naming what filled - where a *session* it has\nno room for is accepted and told in its `CONNACK` that it ends with this\nconnection. The difference is what MQTT lets a server say: a session ending\nwith its connection has a property for saying so, and there is none for \"your\nWill is not held\", so a device that armed one and was accepted would go on\nbelieving the broker would speak for it if it died. It is counted in\n`saguin_connections_refused_total{reason=\"session store full\"}` (RFC 0005),\nand not as a dropped session: it was never accepted.\n\n**A provider that fails for another reason refuses rather than guessing**,\n`0x83`, counted in `saguin_storage_errors_total` and logged: a Will it could\nnot write is not held any more than one it had no room for, and a resumed\nsession it could not read cannot be written back without the subscriptions\nit could not read. The client connects again. **An `UNSUBSCRIBE` it cannot\nstore is refused the same way**, `0x83` for each filter, and the client stays\nsubscribed to them: the `UNSUBACK` is sent once the session is stored without\nthem, since a subscription the client was told was gone would otherwise be\nback after a crash (invariant 18). A 3.1.1 `UNSUBACK` has no code to say it\nwith, so that client is disconnected instead.\n\nThere is no `max_bytes` here, for the reason `broker.retained` has none: the\n**provider's** `max_bytes` bounds the store, and what bounds one session is\n`limits.session_queue_bytes`. Taking a persistent session away from a client\nis the `acl_file`'s `persistent` denial.\n\n**What a session counts against its provider is the memory the broker\nholds for it**, because that is what a session costs while its client is\naway, for up to `limits.max_session_expiry`: 3,500 bytes and its client\nid's length; for each filter 800 bytes, 400 more for a shared one, 600 for\neach level after its first, and the filter's length; and its Will's topic,\npayload and properties. A level two sessions' filters share is charged to\neach, so this is a bound on what they take rather than a measurement of it.\nAt a memory provider's 64MiB that is about 12,000 sessions of one\nthree-level filter each, or about 2,250 of ten four-level filters. A sqlite\nprovider's `max_bytes` bounds its file rather than this memory (\"Bounds on\nwhat a channel holds\"), and `limits.max_session_expiry` says what bounds it\nthere.\n\n**Sessions that fill a memory provider are refused room like anything\nelse, and they leave only as they expire.** A snapshot restored over the\nprovider's `max_bytes` - one written under a larger bound - is loaded whole\nand nothing in it is dropped. Then, until there is room, the provider\nrefuses every publish into a channel it holds (`0x97` where the publish has\nan acknowledgement to carry it), every `SUBSCRIBE` (`0x97`), and every\nresume that carries a Will (`0x97` at `CONNACK`); a broadcast its sessions\nare owed is not kept for them, counted as `storage_full`. A resume without a\nWill is not refused. Retention frees none of that room, because it removes\nrecords and not sessions; the sessions go when they expire, which\n`limits.max_session_expiry` can put a month away. The start says so once,\nnaming the provider and how much of it the sessions hold. The remedy is a\nlarger `max_bytes`, or a sqlite provider for `broker.session.storage`.\n\n**A full session store answers in MQTT's own signals, and keeps nothing by\nhalf.**\n\n| When the store has no room for | The client is told | What is kept |\n|---|---|---|\n| A session asking to outlive its connection, at `CONNECT` | Accepted, with Session Expiry Interval 0 in its `CONNACK` (3.1.1: the session is made clean) | Nothing: the session ends with the connection. Counted in `saguin_sessions_dropped_total{cause=\"storage_full\"}` |\n| A `SUBSCRIBE`, from any client | `0x97` (Quota exceeded) for every filter the packet would add or change (3.1.1: `0x80`) | The session as it was: an existing subscription the packet tried to change goes on being served as it was granted |\n| A QoS 1 or QoS 2 broadcast to sessions that outlive their connections | Nothing: the loss is the sessions', not the publisher's, which is answered `0x00` (RFC 0003 \"Broadcast\") | The broadcast log's oldest messages go to make room, the fewest that make it, and each session is counted once for each message it was owed and lost, in `saguin_session_deliveries_dropped_total{cause=\"storage_full\"}`. A message a session has been sent and has not acknowledged stays, for every session owed it: it is that session's to finish. Where nothing in the log can go, the new message is not kept, and is counted the same way for each session it was for. Both are logged too, as a warning at most every ten seconds with the totals since the last - messages, sessions, deliveries - and the remedy: a provider of its own for `broker.session.storage`, or a larger `max_bytes` |\n| A QoS 1 or QoS 2 broadcast a shared group is owed, or a channel's record a group over the channel is owed | Nothing: the group's loss is not the publisher's, which is answered `0x00` (RFC 0003 \"Broadcast\") | The group's backlog is in the same log, so the row above is what happens: when the log's oldest part goes, what the group was owed from it is counted in `saguin_shares_dropped_total{cause=\"storage_full\"}` |\n\n**The broadcast log gives way to everything else on its provider.**\n`broker.session.storage` is `broker.storage.default` unless it names\nanother, so by default the log shares one `max_bytes` with the channels\nthere, the exactly-once publishes held in them, the retained store and the\nsessions' own records. A write any of them makes that finds the provider\nfull takes its room from the log: the log's oldest messages go as they do\nfor its own next message, counted and logged the same way, the warning\nnaming what the room was for. Only when the log has nothing left it can\ngive is the write refused, answered as its own store answers a full\nprovider. A store's own bound - a channel's `max_bytes` - is not the\nprovider's, and the log gives nothing for it. Channels come first because\nthey are what an operator configures and bounds, where the log holds what\nsessions were owed while they were away; a provider of its own for\n`broker.session.storage` keeps the two apart.\n\n**A `sqlite` provider makes room by the page, not by the message.** Its\nbound is a page count, so what goes frees room only once about a page's\nworth has gone: making room for one message can take several older ones,\nwhere a memory provider gives up exactly the bytes it needs.\n\nA subscription is session state for a client that keeps no session too, so\na full store refuses a new one from every client alike; what is already\nsubscribed goes on matching, at QoS 0 always and at QoS 1 and QoS 2 as the\nrow above allows.\n\n### Validation\n\n**Only the keys this document defines are accepted.** Any other is a\nstartup error that names it, whatever its value, rather than a key the\nbroker ignores while the operator believes it is in force.\n\nThe configuration is validated in full before the broker opens a\nlistener, and every finding is reported, not only the first. An invalid\nconfiguration is a startup failure; there is no partial start in which\nsome channels work.\n\nRules:\n\n- Every channel name satisfies the name rules above.\n- **Every `filter` is a topic filter MQTT would accept**: `#` appears at\n most once and only as the last level, and no level carries a NUL. A\n channel writing no filter is validated as though it had written\n `\u003cname>/#`.\n- **A filter's first level is spelled out and does not begin with ` RFC 0002 - Channels and configuration | Sagüin documentation .** A\n `+` or a `#` there claims `$SYS/…` and `$saguin/…` along with everything\n else, and a channel that swallows the broker's own control topics is a\n broker with no control topics. This is the one restriction on a filter\n that MQTT itself does not make, and it is here rather than in the matcher\n so that the matcher stays the ordinary MQTT one.\n- **No two channels carry the same filter**, compared after `{a,b}` is\n expanded. The error names both channels and the expansion that collided,\n because with braces the two lines in the file need not look alike.\n- **No filter an operator writes carries a level equal to `__dlq`.** Every\n filter that does is one the broker derived for a queue's dead letters.\n- **A queue's filter carries no `{a,b}`.** Its workers pin one exact\n string, and two strings are two consumer groups that each take a copy of\n every job (invariant 4).\n- **Braces hold no `/`**, and neither a brace nor an alternative inside one\n is empty. An alternative is one topic level and never a subtree.\n- **A filter leaves room for a suffix under `limits.max_topic_length`**, so\n that no channel can be configured to claim topics that no publish would\n be allowed to carry.\n- **`broker.storage.providers` defines at least one provider, and\n `broker.storage.default` names one of them.** Sagüin keeps nothing in a\n store the operator did not name, so a configuration without both does\n not start. A memory provider with `snapshot_dir: none` is how to keep\n nothing across a restart.\n- Every `storage` names a defined provider.\n- `retention_bytes` and `max_bytes` are rejected on `latest`, and\n `retention_period` and `retention_bytes` on `queue`, per the table under\n \"Bounds on what a channel holds\" and invariant 2.\n- `dlq_retention_period` and `dlq_retention_bytes` are rejected on any\n channel that is not a `queue`, the only kind that derives one.\n- `broker.retained.storage` names a defined provider, and defaults to\n `broker.storage.default`; `retention_bytes` and `max_bytes` are rejected\n there, per the same table.\n- `default_retention_period` and `default_retention_bytes` are always\n stated. `none` is a value, not an omission: a channel that keeps\n everything says so.\n- A period is a duration and a bytes bound is a size, both under the rules\n below, or the word `none`. Zero is refused for either, and the reason is\n that it has no honest meaning here. Read as a period it says \"keep\n nothing\", which is not a channel; read as Sagüin stores it internally it\n is indistinguishable from `none`, which says keep everything. A value\n that can be argued into meaning either of two opposite things is a value\n to refuse rather than to define, so `none` is the only way to say\n \"keep everything\" and there is no way at all to say \"keep nothing\".\n- A provider's `max_bytes` is at least twice what it holds back for the\n operations which free it - `max_message_size` plus `max_header_bytes`,\n around a megabyte at the defaults; RFC 0003 has why the reserve exists\n and RFC 0004 how each provider keeps it. The error names both figures\n and says which limit to lower.\n- **A `sqlite` provider's `max_bytes` is also above what its empty database\n takes**, 88KiB, plus twice that reserve. The file's tables take pages\n before any record does, and SQLite will not set a ceiling below the pages\n in use, so a smaller figure opens a provider over its own bound that can\n never hold a record.\n- `broker.session.storage` names a defined provider, and defaults to\n `broker.storage.default`. `max_bytes` is rejected there: the provider\n bounds the store.\n- `broker.share.expires_after` parses as a duration in whole seconds.\n\n A *channel's* `max_bytes` takes no such rule. The dead-letter move does\n not check a channel bound at all - a queue and the channel it\n dead-letters into are two different channels, so there is nothing there\n for a reserve to protect.\n- `visibility_timeout`, `job_expires_after` and `retry` are rejected on any\n channel that is not a `queue`.\n- `visibility_timeout` is greater than zero and less than `job_expires_after`,\n and is `30s` where none is written.\n- `retry.max_attempts` is at least 1, and 5 where none is written.\n- `job_expires_after` has no default and may be omitted: a queue that does\n not write it never expires work. It is the only clock that ends a job - a\n publisher's own Message Expiry Interval does not, and reaches the worker\n as `saguin-expires` to decide on instead (RFC 0003 \"`queue` -\n Dead-lettering\").\n- `retry.backoff` is `none`, `linear` or `exponential`, and defaults to\n `none`. `retry.backoff_base` is a duration in whole seconds like every\n other interval here.\n- **`retry.backoff` and `retry.backoff_base` are written together or\n neither is written.** A shape with no base has no gap to build from and a\n base with no shape is read by nothing, so either alone is refused at\n startup rather than starting a broker that looks configured.\n- A sqlite provider declares the `file_path` of its database, and a\n memory provider declares either a `snapshot_dir` or\n `snapshot_dir: none`. Neither takes the other's key: they name\n different kinds of thing and promise different amounts, and accepting\n the wrong one would leave an operator believing they had configured\n something. There is no default, because both defaults are\n wrong: defaulting to a path invents a filename in the operator's\n filesystem, and defaulting to none makes \"durable channel, memory\n provider\" silently mean \"discarded on shutdown\" - the durability\n surprise. The operator states which they meant.\n- A provider no channel names is inert - nothing is created for it, and\n the broker says so once at startup. Declaring one is not an error:\n staging a migration needs the destination provider named before\n anything writes to it (RFC 0004 has the sequence).\n- One broker at a time writes a storage provider. A sqlite provider is\n held by `\u003cfile_path>.lock`, a memory provider by `saguin.lock` in its\n `snapshot_dir`, and a broker that finds one held refuses to start, naming\n the provider and the lock. Two configurations naming one provider is the\n way this happens, since the listener only catches brokers that also want\n the same port.\n\n Both would fail without it, and the memory one fails worse: two brokers\n on one database collide on the primary key - loud, and nothing lost -\n where two on one snapshot directory silently replace each other's\n channels at shutdown and reuse offsets, invariant 9's failure.\n\n The lock is a file of its own so that a sqlite database stays readable\n while the broker runs: `sqlite3` can query a live channel, which locking\n the database itself would prevent. The kernel releases the lock when a\n process ends, so a broker that was killed leaves nothing to clean up -\n but deleting a lock file underneath a running broker does defeat it.\n- One broker at a time serves a Unix socket path, the MQTT one and the\n operations one alike. The socket is held by `\u003cpath>.lock` for as long as\n the broker runs, and what a starting broker does depends on it:\n\n | At the path | The lock | What happens |\n |---|---|---|\n | nothing | free | the socket is created |\n | a socket | free | it was left behind by a broker that crashed, was killed or lost power, and it is replaced |\n | a socket | held | another broker is serving it, and this one refuses to start, naming the lock |\n | anything that is not a socket | either | refused and left alone: a file or directory there is a mistake in the configuration |\n\n A leftover socket is replaced rather than refused because a broker\n restarting unattended after a power cut must come back without anyone\n removing a file. At shutdown a broker removes its socket before\n releasing the lock, so it only ever removes its own, and the lock file\n is kept for the reason above. An abstract name, beginning with `@`, has\n no file to leave behind and takes no lock.\n- Each listener is a block of its own, so a setting belonging to one\n cannot be written beside another. A listener block with no address, or\n a Unix socket with no path, is a mistake rather than a default.\n- `allow_anonymous` is rejected under `broker.operations.listen.tcp` and\n `broker.operations.listen.unix`, and each door's own `password_file` is\n honoured and absolute - \"The operations listener\" above has both\n arguments.\n- **A configuration naming no listener at all gets `tcp` on `:1883`**,\n every interface - reachable by anybody who can route to the box, and\n with no `password_file` it admits them. One listener named is enough to\n turn the default off: a configuration naming just a Unix socket must\n not silently also open the TCP port its author was avoiding.\n- A key naming a filesystem location says which kind it is: `file_path`\n names a file, `snapshot_dir` names a directory. They are not the same\n kind of thing, and one name for both would imply one guarantee.\n- Durations are written as a whole number and one unit, from `ms`, `s`,\n `m`, `h` and `d` - `90m` rather than `1h30m`.\n\n **`ms` is accepted by the form and refused by almost every key that\n uses it.** Nearly every duration in this file is held in whole seconds,\n and the five places a sub-second value is honoured are a bridge's\n `ack_interval`, which is a handful of milliseconds by default; a `sqlite`\n provider's `publish_commit_interval`, which is only useful at that scale,\n and its `flush_interval`, which is 150 milliseconds by default;\n `broker.session.ack_commit_interval`, which is 200 milliseconds by\n default; and `limits.write_timeout`, where a deadline of a few hundred\n milliseconds is a reasonable thing to want on a local network.\n\n What the refusal prevents: a sub-second duration divided into whole\n seconds is **zero**, and zero is a sentinel - `retention_period: 500ms`\n would mean what `none` means, and `visibility_timeout: 500ms` a\n redelivery loop. The refusal is the *key's* - \"held in whole seconds\" -\n not the form's, so an operator is not sent looking for a typo they did\n not make.\n- `snapshot_dir` is a directory, and it is absolute. A memory provider\n writes one file per channel - a queue and its dead-letter channel share\n one, so the move between them cannot be recorded by half - plus a\n manifest naming what was written and when. One file per channel is what\n makes a channel dropped from the configuration a file nobody opens\n rather than a startup error. A relative path is refused, because a\n snapshot must not depend on the directory the broker was started from.\n- A channel name fits a snapshot file name. The encoding percent-encodes\n a dot - so no name can reach `.` or `..` however it is spelled - which\n trebles it: a hundred dots is a legal channel name and a 309-character\n file. It is refused here, where it costs a restart, rather than at the\n shutdown that would lose the channel.\n\nBridges, where a configuration has any:\n\n- A bridge names a `peer`, whose scheme is one of `tcp`, `tls`, `ws`\n or `wss` and which carries a host and a port. `mqtt://` is not a scheme\n anything can dial, and it is refused here rather than at the first\n connection.\n- A bridge names a `client_id`, and there is no default. It is Sagüin's\n identity at the far end - it decides which session the peer resumes,\n and once there is authentication, which principal it is. An invented one\n means two edge boxes silently share a session and take half each other's\n messages.\n- **No two bridges share a `client_id` at one `peer`.** MQTT closes the\n older session when a new one arrives with the same identifier, so two\n such bridges disconnect each other as fast as they can reconnect and\n neither delivers anything, with nothing to show for it but a connection\n log that scrolls. The same identifier at two *different* peers is\n fine: the sessions are at different brokers and never meet.\n- A bridge carries something. That is at least one rule under `topics`. A\n bridge carrying nothing connects, asks for nothing, and holds a session\n open at somebody else's broker for ever.\n- `session_expiry`, `receive_maximum` and `ack_interval` are optional and\n each has a default. Where one is written:\n - `session_expiry` is a duration, no longer than MQTT's own ceiling of\n 136 years, and it is **not** `none`, which every other duration in\n this file accepts: a session that never expires sits on somebody\n else's broker after the edge box is decommissioned, and only that\n operator can clear it. A bridge is a guest, and `none` is the one\n value that makes it a permanent one.\n - `receive_maximum` is between 1 and 65535, which is MQTT's own range for\n it. Zero is not \"unlimited\", it is a subscriber that may receive\n nothing, and MQTT forbids sending it.\n - `ack_interval` is a duration between `1ms` and `1s`. It is one of the\n five keys in this file that may be given a sub-second value, the others\n being a `sqlite` provider's `publish_commit_interval` and\n `flush_interval`, `broker.session.ack_commit_interval` and\n `limits.write_timeout`, and the ceiling is there because the only reason\n to delay an acknowledgement is to batch the ones behind it - which gains\n nothing after a few milliseconds, and past a second is indistinguishable\n from a bridge that has stopped.\n\n Each is refused with what it costs rather than only what the range is,\n because all three are the sort of number somebody raises to make a\n symptom go away: the ceiling arithmetic for `ack_interval`, what is held\n unacknowledged for `receive_maximum`, and whose storage is being spent\n for `session_expiry`.\n- Every rule names a `filter` and a `direction` - `in`, `out` or `both`,\n with no default: an omitted key deciding whether this broker's records\n leave it is not something you could see in the file.\n- `topic:` is required on an `in` or an `out` rule, and refused on a\n `both`, which is the identity mapping (\"Bridges\").\n- **The topic decides where a record lands**, as it does for every\n publisher: a rule names no channel.\n- A rule's `filter` may not name the reserved `$saguin/` space, in any\n direction: inbound it would carry a queue's deliveries across the link,\n taking leases nothing acknowledges (invariant 6). `$SYS/#` is allowed:\n Sagüin refuses the space it defines, not the character. It is allowed\n `in` with a template; `both` on a filter beginning with ` RFC 0002 - Channels and configuration | Sagüin documentation is refused,\n because it maps a topic to itself and nothing may publish a ` RFC 0002 - Channels and configuration | Sagüin documentation topic\n here.\n- A filter is a well-formed MQTT topic filter: `#` is only ever the last\n level, and `+` and `#` each take a whole level.\n- **Two outbound rules with one filter and one topic are refused**, a\n braced spelling meeting a written one included. They would send every\n record to the same topic at the peer twice, and share the one position\n their filter and topic name.\n- A `{a,b}` level is **expanded into one rule per spelling**, and every\n rule above is asked of what it expanded to. The notation is Sagüin's own\n and the far end is an ordinary broker: a brace left in place goes on the\n wire as one literal level, is granted, and receives nothing - a link\n that is configured, reports no error and carries no records. The\n template is unaffected, because `$1`, `$2` … number the filter's `+`\n levels and a braced level is not one.\n\n `fleet/a+b/#` is **refused**: MQTT-4.7.1-2 makes the single-level\n wildcard a whole level, and a publish carrying `+` or `#` in its topic\n is refused everywhere, so no conforming broker can ever hold a level\n spelled `a+b` - the filter is dead by construction and a bridge built on\n one silently carries nothing. A filter that merely has nothing\n publishing to it yet stays legal: `fleet/vessel-99/#` must keep working\n when the vessel comes back, and no startup check can tell it from one\n that never matches.\n- A `topic` uses only substitutions its filter provides: `$1` … `$N` for\n the `+` levels it has, and `$#` only where it ends in `#`. `$#` appears\n only at the end, as `#` does in a filter, because a tail is any number of\n levels and one in the middle makes the shape of the result depend on how\n deep the peer published.\n- **A filter ending in `#` has a topic that uses `$#`.** A discarded tail\n collapses an unbounded number of the peer's topics onto one local topic. On\n an append channel that is a mess; on a `latest` channel it is\n destructive, because two topics become one value overwriting itself while\n the channel does exactly what it is designed to do.\n\n A discarded `$1` is *not* refused. It collapses finitely many topics and\n is a choice an operator can reasonably make - one bridge per vessel drops\n the vessel id on purpose - so only the unbounded case is an error.\n\n **`--check-config` names each rule that discards one**, saying which\n substitution went unused and what it costs, and says separately when the\n channel is `latest`, where the collapse overwrites a value rather than\n crowding a topic. It is a note rather than a finding: the exit code does\n not move and the configuration is still reported `ok`.\n- A `topic` holds no `+` or `#` of its own, does not begin or end with `/`,\n and has no bare ` RFC 0002 - Channels and configuration | Sagüin documentation . A published topic carries no wildcard, a suffix does\n not lead with a separator, and `$x` is a typo that would otherwise become\n literal text.\n\n`--config` names the file and is required: a broker that read whatever\n`saguin.yaml` was in the working directory would serve a different\nconfiguration depending on where it was started from.\n\n`--check-config` performs the whole of the above, touches no storage,\nopens no socket, and exits non-zero on any finding. That exit code is what\nmakes it usable as a systemd `ExecStartPre` and in CI: a bad configuration\nfails the unit rather than the broker. On success it prints the file it\nread and the channels it found, to standard output, so redirect it where\nsilence is wanted.\n\n**It also opens every file the configuration names** - password files,\ncertificates and keys, client and bridge authorities - and reports all of\nthem at once: a schema that parses is not a configuration that starts,\nand a check that read none of those files would pass before the unit died\nanyway, at a restart. Startup reads them the same way, one message naming\nevery unreadable file.\n\nIt still only reads. Nothing is created, nothing is written, no socket is\nopened, so this runs against a configuration from a backup and while a\nbroker is up - but on a machine that does not have the files, it says so,\nwhich is true rather than convenient. `--output` still prints its\ndocument there, because somebody validating a configuration from elsewhere\nwanted the document; the findings go to standard error and the exit code is\nnon-zero.\n\n`--passwd` manages a password file and exits, in the broker's own binary\nbecause Sagüin ships as one file and \"install this other tool to add a user\"\nis not something a single binary gets to say:\n\n```sh\nsaguin --passwd list /etc/saguin/clients.passwd\nsaguin --passwd add /etc/saguin/clients.passwd device-7 [password]\nsaguin --passwd delete /etc/saguin/clients.passwd device-7\n```\n\nThe verbs are `mosquitto_passwd`'s, so an operator who has managed a\nMosquitto fleet already knows them; `list` is an addition, because a\nhashed file cannot be read by eye, and `scope`, which narrows an operator\nto some routes, is the other (RFC 0005 \"Which routes a credential\nreaches\"). With no password on the command line it is prompted for\ntwice, with the terminal's echo off. Like a migration, it reads and\nwrites the file it is given and consults no configuration - which is\nwhat lets a file be prepared before there is a broker to run it, and\nlets one be repaired when the configuration is the thing that is wrong.\n\nNew entries are written `$7 RFC 0002 - Channels and configuration | Sagüin documentation with 1000 iterations, rather than a\nmodern-looking number: the file is read on every CONNECT, and a fleet\nreconnecting after a link drop pays that cost at once. `BenchmarkVerify`\nputs one check at 0.33ms on a Ryzen 7 260 and 0.27ms on a MacBook Pro M1\nPro, and the cost is linear in the count, so a hundred times the\niterations is a hundred times that on every CONNECT. What it costs is\nresistance to offline cracking of a stolen file, and the answer there is\nthe file's permissions - a file Sagüin creates is `0600`, and one that\nexists keeps its mode and owner, as `mosquitto_passwd` keeps them.\n\n`--licenses` prints the licences of the code inside the binary and exits,\nbefore any configuration is read: somebody asking whose code they are\nrunning should not need a configuration file to find out. Sagüin ships as\none file, so there is no `go.mod` beside it and this is the only place that\nanswer exists. The text is generated from the packages that actually reach\nthe binary rather than from everything the module file names, and it\ntravels inside the binary rather than beside it.\n\n**`--output` prints the configuration as it actually resolved** - every\n`!include` expanded into one document, every default filled in, and the\nfile each channel came from written beside it as a YAML comment.\n\n```sh\nsaguin --check-config saguin.yaml # exit 0, prints a summary\nsaguin --check-config saguin.yaml --output # the resolved file, to stdout\n```\n\nIt goes to stdout while the summary and any findings go to stderr, so a\nredirect produces a clean document whatever else was said. What it buys is\nthe one question a configuration split across a dozen files cannot\notherwise answer - what the broker will actually serve - which is worth\nhaving before a start rather than after one. Nothing is created and nothing\nis opened, so it runs while a broker is up, on another machine, and against\na configuration from a backup. `--output` is given with `--check-config`\nand refused without it: a broker that started and also wrote its\nconfiguration to stdout would be writing into whatever its unit's output\nhappens to be.\n\n**What comes out loads again**, and everything else follows from that: a\nchannel that named no provider says which one it got, a queue's derived\ndead-letter channel is left out (a document carrying it would not start),\nand a key nobody wrote is not printed - absence takes the broker-wide\ndefault where the literal `none` is a value and stays. The provenance is\na comment because a reader wants it and a parser must not.\n\n**Credentials are printed as they stand** - no key in the schema holds\none, authentication being configured as paths - and the rule is stated\nfor the first key that does: a document with its secrets starred out\nwould not load. The output is exactly as sensitive as the files it\nflattens, and belongs in the same place with the same permissions.\n"}