From the source of truth

RFC 0005 - Operations

A static snapshot from the Sagüin repository.

-prefixed base64, whose alphabet has none, so the\nsecond colon ends the hash and no realistic value is misread. A delimiter\ninside the name would not survive `alice@example.com`, which this file\naccepts.\n\n**The routes are paths rather than names for groups of them**, so there is\nno mapping in the code that an operator cannot read. The price, said rather\nthan discovered: if a path ever changes, the files change with it.\n\n**Matched at a segment boundary.** `/v1/operations` grants\n`/v1/operations/consumers` and does not grant `/v1/operationsomething` - a\nroute nobody wrote, reached by a name that merely starts the same way.\n\n**A scope naming no route Sagüin serves is a startup error**, and\n`--check-config` reports it: it denies that user everything while reading,\nfrom the file, as though it granted something.\n\n**A third field on an MQTT password file is a startup error too.** An MQTT\nclient reaches topics rather than paths, its permissions are the\n`acl_file`'s, and a scope there would govern nothing. It is also what keeps\nthat file Mosquitto's format hash for hash, which RFC 0002 promises anybody\nmigrating.\n\n**A credential that is good but does not reach a route is answered `403`,\nnot `401`.** A `401` says \"try again with a credential\", so a scraper\nretries and records an authentication failure - sending an operator to check\na password that is correct, when the answer is that this credential does not\nreach here.\n\n`saguin --passwd scope \u003cfile> \u003cuser> \u003croutes|all>` writes the field. It sets\nthe whole list rather than adding to it: a user with no routes named reaches\nevery one, so an \"add\" against such a user would have to narrow it, which is\nan add that removes.\n\n**A `/metrics` with no credential is a loopback-only `/metrics`** - the\nrule below refuses to start on a routable address until something there\nauthenticates. **The credential applies to every transport the listener\nanswers on**, the Unix socket included: its permissions are a second gate\nrather than an alternative, so a local agent reading through the socket\nneeds the credential from the day one is configured.\n\n**Two contracts Sagüin conforms to, and one it owns.**\n\n`/health` answers an orchestrator's probe and `/metrics` answers a\nscraper. Neither shape is Sagüin's to choose - one is a liveness\nconvention, the other Prometheus's exposition format - so neither carries\na version and neither ever changes. `/v1/…` is Sagüin's own, will grow,\nand will one day change, which is why it says which version it is: a\n`/v2` can stand beside a `/v1` while people move.\n\n**No route writes anything, no user interface, and no payloads - ever.\nAnd no secret: not a password, not a hash, not the contents of a key\nfile.** That is the rule, and it is bounded on what comes out rather than\non how coarse it is. The records themselves come out the way every other\nclient reads them: through MQTT, held to the same permissions as everybody\nelse.\n\n**No secret can reach these routes, and that is the schema's doing\nrather than a filter's.** Nothing in the configuration holds a\ncredential - every one is a path - so there is no redaction pass to keep\nin step with the schema, and a key that carried a secret would be a\nchange to what Sagüin stores, refused on those grounds. What holds it\nthere is a test that reads every file the response names and requires\nthat none of its contents came back. The promise is bounded on what\ncomes out: nothing that leaves this listener is a credential, and\nnothing that goes into it changes anything.\n\n**The `/v1` routes are not metrics and must not be scraped.** They answer\nwith a list rather than a number, on demand rather than on a timer, and\nthey carry exactly what the catalogue is closed against - a client id is a\nstring a client chose, so a series keyed by one is a series count chosen by\nwhoever connects. A scraper pointed at them writes an unbounded number of\nseries into somebody's monitoring system, which then falls over, some\ndistance from Sagüin and long after the cause.\n\n**Which is why they exist as a route rather than as metrics.** \"Which of my\nthree hundred devices is behind\" and \"what is actually stuck in that queue\"\nare the two questions an operator arrives with, and the catalogue can answer\nneither: `saguin_channel_consumer_position_min` is one number per channel,\nso one straggler reads exactly like a fleet that has stopped, and\n`saguin_queue_depth` says four hundred while nothing says what. Both answers\nare in the store already. The only other way to see a queue's contents is to\nconsume them, which takes work from the workers the viewer was sent to\ndiagnose.\n\n**Every row is capped, and the cap is a hundred.** Invariant 13 is about\neverything that accumulates and a response body accumulates: without a cap,\na route that lists positions can be asked to read a fleet of ten thousand\ninto memory and out over a socket. **Worst-first ordering is what makes the\ncap safe** - the rows that fall off the end are the ones nobody was looking\nfor - and every response carries how many exist beside how many it returned,\nbecause a caller shown a hundred of three hundred and not told so believes\nit has seen the fleet.\n\n**A queue that is not configured answers 404 and does not name the ones that\nare.** These routes are behind a credential, but a 404 that lists the\nchannels is a route that returns the configuration.\n\n**JSON, which is the first thing on this listener that is not\n`text/plain`.** `/metrics` is Prometheus text because a scraper reads it;\nthese are read by a person and by `jq`.\n\n**`/health` needs no credential, and that is the point of it** - a\nliveness probe that authenticates is a probe that restarts a healthy\nbroker over a secret; \"The health endpoint\" below has the whole of it.\n\n**No block means no listener.** A `broker.operations` that is absent opens\nnothing, exactly as a `listen` with no `unix` block opens no socket. The\naddress is bound at startup with the MQTT listeners, so a port already in\nuse is a startup error naming it rather than a surprise later.\n\n**The block is named for what every path on it has in common**, which is\nthat each is a read-only fact for the person running the broker rather\nthan anything a client can use - and it is this document's own title, so\nthe file and the specification name the same thing the same way.\n\n## Authentication and TLS\n\nWhat is written below is the shape of each, and the one rule that stops\nauthentication being deferred for ever.\n\n**Authentication is required on `/metrics`**, as HTTP Basic in the request\nheader. An operator reaches Sagüin from somewhere else on the network, the\nsame way an MQTT client does, so this cannot be a localhost-only interface\nwith the question deferred. Basic is what every scraper already sends and\nwhat every proxy already understands; it carries the credential in clear,\nwhich is what the paragraph on TLS below is about.\n\nThe operators are `broker.operations.password_file` - or a listener's\nown, which wins for that door alone - in Mosquitto's format, managed with\n`saguin --passwd`, read at startup and re-read on `SIGUSR1` (RFC 0002\n\"Who may connect\"). A file naming no users is a startup error rather than\na listener nobody can read.\n\n**The credentials are not the MQTT credentials.** An operator is not a\ndevice. One credential set granting both topic access and broker\nstatistics means every device that can publish can also enumerate the\nchannels, their volumes, their consumers' positions and the address of\nevery broker this one dials - and a fleet's credentials are on the fleet,\nwhere they are readable by anyone holding one device.\n\n**TLS is optional**, and configured as `broker.operations.listen.tcp.tls`\nwith a certificate and its key (RFC 0002 \"TLS on a listener\"). On an edge\nbox the operations port is commonly reached across a management network or\nan SSH tunnel that already carries the transport security, and requiring a\ncertificate there produces a self-signed one that nobody rotates and every\nclient is told to ignore - which is worse than plain HTTP on a trusted\nlink, because it looks secured. Where the port crosses a network the\noperator does not own, TLS. The Unix socket takes none: it does not leave\nthe machine.\n\n### Who Sagüin thinks you are\n\n**Five ways to be named, and the last two are believed only at a door the\noperator chose who may open.**\n\n| Named by | Where it counts |\n|---|---|\n| A password this broker checked | anywhere |\n| A client certificate an authority named in `tls.client_ca_file` signed | anywhere |\n| A certificate that carries no name at all | anywhere, and it is *not* nobody - see below |\n| `X-Saguin-Principal`, set by a reverse proxy that did the checking | a Unix socket, and nowhere else |\n| PROXY protocol v2, carrying the Common Name the proxy verified | a Unix socket with `proxy_protocol: true` |\n\nThe header is whatever the caller typed. On a routable address anybody who\ncan reach the port can write it, so it is not read there at all rather than\nread and weighed - a bypass that needs a rule to be safe is a bypass.\n\n**A Unix socket and a loopback port are not equally strong doors, and only\nthe socket carries a name.** A socket has file permissions: the operator\ndecides which user or group may speak to it. A loopback port has none -\nunreachable from the network is not the same as reachable only by whoever\nthe operator chose, because *every* process on the machine may connect. A\nheader there would let any local process name itself any operator and read\nwhich devices are behind, without the password the operators file exists to\ncheck, so it is not read there.\n\n**This is the rule the MQTT listener's `proxy_protocol` already\nstates** (RFC 0002): a socket's file permissions decide who may assert a\nname, and a TCP listener would need a trusted-proxy allowlist first.\n\nSo a reverse proxy that names its caller reaches Sagüin through the socket:\n`proxy_pass http://unix:/run/saguin/operations.sock:/v1/` for the header -\nthe trailing `/v1/` matters, because a `proxy_pass` carrying a URI\nreplaces the matched location prefix and a bare `/` would strip it - or\nthe stream module with `proxy_protocol v2` for the Common Name. A loopback\nport is still the right door for a scraper with a password, or for a local\nreader with none - it just cannot carry somebody else's word for who is\nasking.\n\n**A name with no entry in the password file reaches `/metrics` and nothing\nelse.** Being scraped needs no entry, which is what a certificate or a\nproxy is almost always arranged for; reading which device is behind names\ndevices, and that is a thing an operator says out loud by putting the name\nin the file. Widening it is `saguin --passwd scope`, the same answer as for\neverybody else.\n\n**A name is the Common Name, or the first DNS name when there is none.**\nCN is where an operator's own authority usually puts it and what an\nacl_file is written against, but it has been deprecated as an identifier\nfor years and many authorities now issue a subject alternative name and\nnothing else. **A name is taken exactly as the certificate, the proxy or\nthe header states it**, spaces and all, as the MQTT listener takes a\ncertificate's name: one certificate is one identity at either door. A\nheader has no edge spaces to keep: HTTP does not count them as part of its\nvalue. **A certificate this broker verified and cannot name is still not\nnobody**: it came through a door that admits nobody without one, so it\ngets what a named stranger gets - `/metrics` and no more. **A name holding\nU+0000 or a control character is nobody's**, from any of the four sources,\nand the request is answered `403` with a line in the log, as the MQTT\nlistener refuses the same name (RFC 0002 \"TLS on a listener\").\n\n**A credential that was offered and failed ends there.** A wrong password\nis a refusal, not a fall-through to whatever else might name the caller -\notherwise a stale password beside a proxy's header answers 200, which is a\nrejected credential that was not rejected.\n\n**Nobody is not the same as somebody unscoped.** A request that carries no\nname at all - no password, no certificate, no proxy - is trusted by the\ndoor it arrived at, which validation has already made a loopback address or\na socket. That is the local reader with no credential to give, and it\nreaches everything.\n\n**And that local reader answers to an address, not to a name.** Unreachable\nfrom the network is not unreachable from the internet. A page on another\nsite whose domain re-resolves to `127.0.0.1` is *same-origin* to this\nlistener and reads it through whatever browser has it open - the catalogue,\nthe sessions, the ACL, the user list, the resolved configuration. Every\nroute here is a `GET`, so nothing can be changed that way; all of it can be\nread, and the operator's own browser is the one asking. The one thing that\npage cannot forge is the `Host` header, because a browser sends the name it\nwas loaded by. So a request carrying no name at all is served when it\naddressed this broker - an IP literal, or `localhost` - and refused with\n`403` naming the setting that lifts it when it asked for this broker by a\nname.\n\n**Only there.** A listener with a password file, a client certificate or a\nproxy that names its caller never reaches that branch, so a scraper or a\nviewer calling a real deployment by its hostname is untouched - and by the\nrule above, anything off this machine has one of those. A Unix socket is\nexempt: no browser can open one, and whatever proxy is in front of it sends\nthe name it was asked for. `/health` is not behind it either, being never\nauthenticated and carrying nothing about this broker's contents, precisely\nso that a probe which cannot hold a credential can read it from wherever it\nruns.\n\n**A required client certificate takes `/health` with it, and that is the\none thing to arrange around.** A liveness probe which authenticates is a\nprobe that fails when the credential is wrong, expired or not yet mounted -\nso a broker in perfect health gets restarted. Mutual TLS on the only TCP\nlistener produces exactly that: the probe has no certificate, the handshake\nends, and the probe sees a connection error rather than a 200. Give the\nprobe its own door - the Unix socket, or a second loopback listener - rather\nthan a certificate.\n\n**And a required client certificate lifts the loopback rule**, for the\nreason the socket already lifts it: nothing without a certificate this\nbroker's authority signed completes the handshake, so the port is not\nreachable by whoever can route to it. `require_certificate: false` does not\nlift it - that is the mixed mode where a client presenting nothing still\nconnects.\n\n**Behind a reverse proxy**, where a deployment already terminates TLS\ncentrally, the broker's own listeners stay on a socket and loopback and\nthe proxy carries the credential and the certificate. What follows is\nthe whole working pair. The stream half is the MQTT listener's\narrangement, specified in RFC 0002 under \"Behind a proxy that terminated\nTLS\", and is shown here so the pair is one file nginx can load:\n\n```nginx\n# Raw MQTT. proxy_protocol v2 is what puts the client's address and its\n# certificate's Common Name back in the broker's hands - `proxy_protocol\n# on` is v1, carries no TLVs, and saguin refuses it by name.\nstream {\n server {\n listen 8883 ssl;\n ssl_certificate /etc/saguin/tls/cert.pem;\n ssl_certificate_key /etc/saguin/tls/key.pem;\n ssl_client_certificate /etc/saguin/tls/clients-ca.pem;\n ssl_verify_client on;\n\n proxy_pass unix:/run/saguin/saguin.sock;\n proxy_protocol v2;\n\n # Above the largest keepalive in the fleet. See below.\n proxy_timeout 20m;\n }\n}\n\nhttp {\n # The CN of the certificate nginx verified. Stock nginx has no\n # variable for it - $ssl_client_s_dn is the whole subject.\n map $ssl_client_s_dn $client_cn {\n default \"\";\n \"~(^|,)CN=(?\u003ccn>[^,]+)\" $cn;\n }\n\n # MQTT over WebSocket, for browsers.\n server {\n listen 8443 ssl;\n ssl_certificate /etc/saguin/tls/cert.pem;\n ssl_certificate_key /etc/saguin/tls/key.pem;\n location /mqtt {\n proxy_pass http://127.0.0.1:8083;\n proxy_http_version 1.1;\n proxy_set_header Upgrade $http_upgrade;\n proxy_set_header Connection \"upgrade\";\n proxy_read_timeout 3600s;\n }\n }\n\n # /metrics behind a credential. /health is deliberately not here:\n # a liveness probe that goes through the proxy reports on the proxy.\n server {\n listen 9443 ssl;\n ssl_certificate /etc/saguin/tls/cert.pem;\n ssl_certificate_key /etc/saguin/tls/key.pem;\n\n # Without these, nginx asks for no certificate, $client_cn is\n # empty and proxy_set_header sends nothing - so nothing names the\n # caller and saguin answers 401 rather than serving a stranger.\n # `optional` rather than `on`, so /metrics below still works for a\n # scraper that has a password and no certificate.\n ssl_client_certificate /etc/saguin/tls/operators-ca.pem;\n ssl_verify_client optional;\n location /metrics {\n auth_basic \"saguin\";\n auth_basic_user_file /etc/saguin/metrics.htpasswd;\n proxy_pass http://127.0.0.1:9090;\n }\n\n # Or let the client certificate be the credential and tell saguin\n # whose it was. `$ssl_client_s_dn_cn` does not exist in stock\n # nginx; the subject DN does, and a map takes the CN out of it.\n # Through the socket, because that is the only door that carries a\n # name: its file permissions decide who may claim one.\n #\n # The `:/v1/` on the end is load-bearing. A proxy_pass carrying a\n # URI replaces the matched prefix, so `:/` would send\n # /v1/operations/consumers on as /operations/consumers and saguin\n # would answer 404 to every request through here.\n location /v1/ {\n proxy_pass http://unix:/run/saguin/operations.sock:/v1/;\n proxy_set_header X-Saguin-Principal $client_cn;\n }\n }\n}\n```\n\n**What each door answers**: `/metrics` with a password and no\ncertificate, 200; with a wrong one, 401. `/v1` with a certificate the\nauthority signed, 200 when that name has an entry in Sagüin's operators\nfile and 403 when it does not. `/v1` with nothing at all, **401** -\nnginx sends no name, so Sagüin has none and refuses. And a name a\nclient sends itself never survives, because `proxy_set_header` replaces\nit: a caller presenting no certificate and writing\n`X-Saguin-Principal: operator` is answered 401.\n\nThat 401 row is the one to check after any change to this block: a\nblock without `ssl_client_certificate` is syntactically perfect and\nopen, and `nginx -t` cannot see that - `make nginx`, which drives this\nblock against a real nginx and asserts every row above, can.\n\nThe broker beneath it:\n\n```yaml\nbroker:\n mqtt:\n listen:\n unix:\n path: /run/saguin/saguin.sock\n mode: \"0660\"\n proxy_protocol: true\n ws:\n address: 127.0.0.1:8083\n operations:\n listen:\n tcp:\n address: 127.0.0.1:9090 # /metrics, behind nginx's auth_basic\n unix:\n path: /run/saguin/operations.sock # /v1, where a name can be carried\n mode: \"0660\" # nginx's worker needs this group\n password_file: /etc/saguin/operations.passwd # and this is not optional\n```\n\n**`/health` is deliberately not proxied.** A liveness probe routed through\nthe proxy reports on the proxy: it answers while the broker is gone, and\nstops answering when the proxy restarts under a broker that is fine. It is\nreached on loopback, which is also why it carries no credential - the\nparagraph on the two endpoints above says why that asymmetry is the rule\nrather than an oversight.\n\n**That `password_file` is what makes the arrangement fail closed, and\nleaving it out is not a smaller version of it.** The proxy's credential\nis on top of Sagüin's rather than instead: without the file, Sagüin has\nnobody to refuse - a request carrying no name is read as a local reader\nwith no credential to give, which reaches every route. With `optional`\nverification above, a caller presenting no certificate is exactly that:\nnginx sends an empty header, and the door opens. Without the file,\n`/v1` with no certificate answers **200**; with it, **401**.\n\n**Note the two files.** `auth_basic_user_file` is nginx's own, in\nhtpasswd format - not Sagüin's operators file, which is Mosquitto's\nPBKDF2 format and which nginx answers 500 on. The scraper goes in both,\nbecause nginx passes the `Authorization` header through: a scraper\nnginx authenticated and Sagüin does not know is answered 401 by Sagüin.\nIf one gate is enough for a deployment, it is Sagüin's: drop\n`auth_basic` and let the broker answer.\n\n**The rule that makes \"authentication later\" impossible: Sagüin refuses to\nstart when the operations listener's TCP address is not a loopback address\nand no `password_file` governs it - neither its own nor the block's -\nnaming both.** The failure it prevents is the ordinary one, not a careless\none. The port starts on localhost with no credential configured, which is\ncorrect; then a scraper arrives on another box, somebody changes the\naddress to `0.0.0.0:9090` because that is what makes it reachable, and the\nbroker begins serving channel names, message volumes, consumer positions\nand its bridges' peer addresses to anyone who can route to it - with\nnothing anywhere saying that a decision was made. A startup error is where\nthat gets caught, because it is the one moment the operator is looking.\n\nA rejected credential is `401` with a `WWW-Authenticate` header, and the\nlog line says a credential was rejected and does not contain it.\n\n## The health endpoint\n\n`GET /health` answers `200` and `{\"status\":\"ok\"}`; `HEAD` gets the same\nstatus and no body; another method on it is `405`.\n\n**It answers one question: would restarting this process help?** That is\nthe only question a liveness probe is entitled to ask, and it is why the\nendpoint deliberately reports nothing about the world outside the process.\n\nAn inbound bridge whose far end is down is a real thing an operator wants\nto know and a terrible thing to fail a probe on: the broker would be\nrestarted, repeatedly, because somebody else's broker is unreachable - and\nthe restart cannot fix it. Storage errors, bridge state, and how far behind\neach bridge is are all worth reporting and none of them belongs here. They\nbelong to the metrics below.\n\n**The metric is the state and the log is the narrative**:\n`saguin_bridge_connected` is where an operator reads whether a link is up\nnow, whatever their log level, and a line that retracts another is\nwritten at the level of the line it retracts (RFC 0002 \"How much the\nbroker says\").\n\nWhat is left is not nothing. Sagüin has one broker-wide mutex on the\npublish path, and a broker that cannot take it is one whose publishes have\nstopped - which is the failure a restart does fix. So the handler takes\nthat lock, with a timeout, and `503` and `{\"status\":\"unavailable\"}` means\nit could not.\n\n**No path in Sagüin holds that lock across disk I/O**, which is what makes\nthe answer unambiguous. A snapshot builds its state under the lock and\nreleases it before writing the file; the retention sweep copies the store\nmaps under the lock so that no removal is done holding it. A hold long\nenough to fail this probe is therefore not slow hardware and not a large\nchannel - it is a defect. **That is also a constraint on everything else on\nthis listener**: a scrape that formatted its output while holding the lock\nwould make this endpoint flap on the same port, and the flapping would be\ntelling the truth. The metrics section repeats the rule where it bites.\n\n**The timeout is a constant rather than a key.** An operator has no real\nreason to change it. A publish through the wire costs tens of microseconds\n(\"What it costs, measured\" below), so the lock is held for less than that\nat a time, and any threshold separating \"busy\" from \"wedged\" is four orders\nof magnitude away from both. A knob whose only settings are the default\nand wrong is not a knob.\n\n**There is no `/ready`.** Snapshots are loaded before any listener binds,\nso there is no moment at which Sagüin is running but not ready to serve,\nand a second endpoint that always answers what the first one did is that\nsame non-knob in another shape. If bridges ever have to be connected\nbefore the broker serves, that stops being true, and this is the paragraph\nthat would change.\n\n**The body carries nothing else, and that is a decision rather than an\nomission.** It is the one path on this listener with no authentication, so\nwhatever it returns is readable by anyone who can reach the port: a\nversion, a channel name, a record count, all of it. Adding a field later\nis easy. Taking one back after somebody's monitoring has started parsing\nit is not.\n\n## Metrics\n\n`GET /metrics` answers the catalogue below in the Prometheus text\nexposition format, which every scraper in common use reads and with which\nOpenMetrics is compatible.\n\n**Pulled, never pushed, and that is the resource decision.** Sagüin runs\non a gateway, a Pi, or a vehicle, and its job is MQTT. A broker that\ncomputes its own statistics on a timer pays for them every interval for\never, whether or not a single person is looking; a broker that answers\nwhen asked pays only when somebody asks, and the cadence is set in the\nscraper the operator already runs. A `$SYS` tree is the timer version of\nthe same idea, and the section below says why Sagüin has none.\n\n### What a metric is allowed to cost\n\n**Where the substrate already counts something, that counter is the one\nread.** It maintains connections, subscriptions, messages, packets and\nbytes as atomics on the paths that change them - the numbers its own `$SYS`\ntree publishes, though Sagüin silences that tree. Silencing stops the\npublishing and not the counting, so reading them here costs an atomic load\nand duplicates nothing. Keeping a second tally beside one of those would be\ntwo numbers to keep in step and one of them going quietly wrong, which is\nthe reasoning that refused the `$SYS` tree in the first place.\n\n**A metric is a number the broker already holds, or it does not ship.**\nThis is the rule the catalogue was built against rather than a note about\nthe implementation, because the failure it prevents is one nobody\nattributes correctly: a monitoring tool, added to find out why the broker\nis slow, is what is making it slow.\n\nThree costs, measured against the stores as they are:\n\n**Free - the number is a field.** The offset a channel will assign next,\nthe retention floor, the bytes an `append` or `queue` channel holds, the\ncount of records out with queue workers, and every counter. Both storage\nproviders keep these in memory and advance them from committed\ntransactions.\n\n**Free because the store was asked to keep it, which is not the same as\nfree to ask for.** How many topics a `latest` channel holds a current value\nfor is the size of a map the `memory` store is already keeping. The\n`sqlite` store counts its rows once, when the channel opens, and advances\nthat number from committed transactions - the arrangement a queue's depth\nalready has, and for the same reason: counting rows on a scrape is the one\ncost this document refuses outright.\n\n**What makes it free on the write path is the order of the write.** An\nupsert reports success whichever branch it took, so a store using one\ncannot tell a new topic from a replacement without asking again. Updating\nfirst and inserting only where nothing matched gives that answer as a side\neffect, and costs nothing over an upsert - it goes straight to the row\nrather than attempting an insert and hitting the conflict:\n`BenchmarkLatestSet`'s sqlite row is 18.6us. A topic seen for the first\ntime costs 27.5us, once in that topic's life - ten thousand devices pay\n27ms between them, ever.\n\n**A `latest` channel's bytes are free on neither provider**, for a reason\nthe count does not share. A write replaces a value, so the change in size\nis a difference and measuring it means reading the outgoing value first.\nThe count escapes that because a replacement does not change it: only a\ntopic arriving or leaving does, and both are already known where they\nhappen. The catalogue says where the bytes are absent.\n\n**Proportional to the number of consumers - and this one is paid.** The\nworst lag on a channel means the lowest stored position on it, and there is\none position per durable consumer, so a fleet of ten thousand devices\nreading one channel is ten thousand positions to take a minimum over, per\nchannel, per scrape. The memory store walks its map; the sqlite store asks\nits table for `min(\"offset\")`.\n\n`BenchmarkLowestPosition`, at `-benchtime 2s -count 2`, one channel:\n\n| consumers on one channel | memory - Ryzen 7 260 | sqlite - Ryzen 7 260 |\n|---|---|---|\n| 100 | 531ns | 25.6us |\n| 1,000 | 5.7us | 174us |\n| 10,000 | 59.6us | 1.69ms |\n\n**Those are the cost of the minimum and the count together.** They are one\npass: the memory store reads the size of the map it is already walking, and\nthe sqlite store asks for `min()` and `count()` in one statement over the\nsame rows, where a second method, asking separately, would double it.\n\n**The minimum is computed on every scrape rather than cached**, because a\ncached minimum that only falls never rises when consumers catch up, and\n`floor > position_min` - the one alert below - would then fire on a broker\nwhere every consumer had caught up and nothing had been deleted unread. So\nthe reading above is the price of the alert being usable: **1.69ms, once\nper channel, per scrape**, at a fleet size on the large side of what a\nsingle-node broker is for, against a scrape interval with a configured\nfloor (`min_scrape_interval`).\n\n**Proportional to the number of records - never.** Counting the rows of a\nhistory is a table scan, and a scraper doing it every fifteen seconds\nacross every channel is a load source wearing a monitoring tool's clothes.\nTwo consequences, and both are why the catalogue looks the way it does:\n\n- **A channel's record count is `next − floor`**, which is two of the free\n numbers above. Retention only ever removes a prefix and advances the\n floor past it, so no record is ever missing from the middle of a log and\n the subtraction is exact.\n- **A queue's depth is not derivable**, because resolution removes a\n record from the middle rather than the front. It is the one number in\n the catalogue that a store does not already hold, so the store keeps a\n count in memory the way it already keeps the byte total, and the scrape\n reads that instead of counting rows.\n\n**And the scrape does not hold the publish path.** The numbers are copied\nunder the broker's lock and formatted outside it, which is the discipline\nthe snapshot and the retention sweep already follow. A scrape that\nformatted under the lock would stall publishes and fail the health probe\nsitting on the same listener.\n\n### What it costs, measured\n\nOne scrape, at four configuration sizes. A tenth of the channels are queues\nand a tenth are `latest`, which is the shape of a real file rather than a\nthousand of one kind - the queue series are the widest part of the\ncatalogue, so measuring over append channels alone would understate it.\n`BenchmarkScrape` takes it, every channel holding one record, at\n`-benchtime 2s -count 2` on an AMD Ryzen 7 260 (8 cores, 16 threads):\n\n| Channels | Response | Gzipped | CPU, plain | CPU, gzipped |\n|---|---|---|---|---|\n| 10 | 12.3 KiB | 2.9 KiB | 33 µs | 291 µs |\n| 100 | 55.2 KiB | 6.0 KiB | 131 µs | 736 µs |\n| 1,000 | 493 KiB | 34 KiB | 1.65 ms | 4.06 ms |\n| 10,000 | 4.85 MiB | 291 KiB | 15.4–16.5 ms | 32–33 ms |\n\n**Each series per channel adds to both columns** - roughly 5–13% to the\nbody and 12–35% to the CPU per row of the catalogue - which is the\nstanding cost of adding one, worth knowing before the next is.\n\nA publish through the wire, one publisher waiting for each acknowledgement\n(`BenchmarkPublishThroughMQTT`), costs 30–32µs on a memory channel and\n**92–106µs** on a sqlite one - so a hundred-channel scrape is about four\nmemory publishes, once per scrape interval, and gzip is most of the cost of\na scrape that asks for it.\n\n**Two conditions on the sqlite figure.** `TMPDIR` must be on a real disk:\n`/tmp` is a tmpfs on most machines, so an unset one measures the page cache\nand returns about 78µs, which is not a number any deployment sees. And the\nrow varies by about 15% between runs on one machine and one disk, so take\nseveral counts and quote the range, not the run. The memory row does not do\nthis: it repeats to the microsecond.\n\n**Both the response and the CPU grow with the channel count**, which is\nstated here so that an operator can predict it from the number of channels\nthey configured rather than discover it. RFC 0004 measures the other two\ncosts that grow the same way: the ten thousand snapshot files at a\ngraceful shutdown, and the retention sweep. The sweep is two different\ncosts - its size half reads nothing and is the same at ten thousand\nchannels as at ten, and its age half reads once per channel on a `sqlite`\nprovider, on an interval the shortest configured retention decides. So the\nfigure to plan against is this scrape and that age sweep, and which of\nthem dominates depends on how often each is asked for.\n\nGzip is offered when the client asks for it, which Prometheus does by\ndefault. It returns about a fifteenth of the bytes at a thousand channels,\nwhich is the right trade on any link an operator would scrape across and\nthe wrong one on none.\n\n### The observer does not set the cost\n\n`min_scrape_interval` is the shortest interval at which the catalogue is\nrecomputed. A scrape arriving sooner is answered from the previous one.\n\n**The cost of being observed is bounded by Sagüin rather than by the\nobserver.** Without it, the load is chosen by whoever configures the\nscraper - and by *how many* of them, since five monitoring systems at one\nsecond each is five times the bill, which is not a number this broker got\nto agree to. With it, one interval is one computation however many people\nask and however fast.\n\n**A scrape arriving early is answered, not refused.** Prometheus records a\nrefused scrape as a *failed* one: a gap in the graph and, usually, an\nalert. A defence that turns into the operator's incident is not a defence,\nso the early scrape gets the previous bytes and nobody is told off.\n\nWhat it costs is that a sample can be up to one interval old while the\nscraper stamps it with the time it asked.\n\n**A minute is the floor and the default, and a shorter one is a startup\nerror.** Not raised quietly: an operator who writes `10s` and is silently\ngiven a minute reads their graphs believing they have ten-second\nresolution, and every conclusion they draw about the shape of a spike is\ndrawn at the wrong scale. Naming the floor costs one restart.\n\n**A minute is what a broker's metrics are worth.** The numbers do not change\nusefully faster than that, and reading them costs the broker rather than the\nreader. Against a scraper set to Prometheus's own default of fifteen seconds\nit is that bill cut fourfold, and the bill is measured above: 1.69ms per\nchannel per scrape at ten thousand consumers on the sqlite store.\n\n**It bounds how late the alert below can be**, which is the figure to weigh\nif it is ever raised: retention removes records on a scale of days, so a\nminute is nowhere near it.\n\n**Set the scrape interval at or above this one.** A scraper running faster\nsees the same numbers repeated until the interval passes, and a counter\nthat repeats and then jumps makes a rate calculation look like a staircase.\nIt averages out correctly over any window worth alerting on, and it is\nconfusing to look at, so the two intervals should agree.\n\n### What a scrape looks like\n\nTaken from a running broker with an `append`, a `latest` and a `queue`\nchannel on a memory provider, and a fourth channel, an `append` on a\n`sqlite` provider, that one inbound bridge fills from its peer. A durable\nconsumer is three records behind the head, a worker holds one of three\njobs, one publish went to a topic no channel claims and nobody subscribes\nto, one publish into the dead-letter channel was refused, one device lost\nits link and its Will was published, and one connection was refused for\nnaming an authentication method. `HELP` and `TYPE` precede every name and\nare elided here after the first, and so are the response headers other\nthan the content type and the value of the `version` label, which is\nwhatever the build says it is. Every name in the catalogue has a line but\n`saguin_storage_errors_total`, which has none until a provider fails, and\n`saguin_subscriptions_refused_total`, which has none until a filter is\nrefused:\n\n```\n$ curl -s -D - http://127.0.0.1:9090/metrics\n\nHTTP/1.1 200 OK\nContent-Type: text/plain; version=0.0.4; charset=utf-8\n\n# HELP saguin_build_info Always 1. This is where saguin's own version is stated.\n# TYPE saguin_build_info gauge\nsaguin_build_info{version=\"…\",broker_id=\"edge-1\"} 1\nsaguin_uptime_seconds 9.613964048\nsaguin_connections 1\nsaguin_connections_by_protocol{protocol=\"5\"} 1\nsaguin_subscriptions 2\nsaguin_max_session_expiry_seconds 2592000\nsaguin_channel_info{channel=\"events\",filter=\"iot/+/events/+\",type=\"append\",provider=\"mem\"} 1\nsaguin_channel_info{channel=\"jobs\",filter=\"iot/+/work/+\",type=\"queue\",provider=\"mem\"} 1\nsaguin_channel_info{channel=\"jobs__dlq\",filter=\"iot/+/work/+/__dlq\",type=\"append\",provider=\"mem\"} 1\nsaguin_channel_info{channel=\"readings\",filter=\"iot/+/readings/+\",type=\"append\",provider=\"disk\"} 1\nsaguin_channel_info{channel=\"state\",filter=\"iot/+/state/+\",type=\"latest\",provider=\"mem\"} 1\nsaguin_channel_bytes{channel=\"events\"} 390\nsaguin_channel_bytes{channel=\"jobs\"} 195\nsaguin_channel_bytes{channel=\"jobs__dlq\"} 0\nsaguin_channel_bytes{channel=\"readings\"} 126\nsaguin_channel_next_offset{channel=\"events\"} 7\nsaguin_channel_next_offset{channel=\"jobs__dlq\"} 1\nsaguin_channel_next_offset{channel=\"readings\"} 3\nsaguin_channel_floor_offset{channel=\"events\"} 1\nsaguin_channel_floor_offset{channel=\"jobs__dlq\"} 1\nsaguin_channel_floor_offset{channel=\"readings\"} 1\nsaguin_channel_records{channel=\"events\"} 6\nsaguin_channel_records{channel=\"jobs__dlq\"} 0\nsaguin_channel_records{channel=\"readings\"} 2\nsaguin_channel_records{channel=\"state\"} 2\nsaguin_channel_consumer_position_min{channel=\"events\"} 4\nsaguin_channel_consumers{channel=\"events\"} 1\nsaguin_channel_consumers{channel=\"jobs__dlq\"} 0\nsaguin_channel_consumers{channel=\"readings\"} 0\nsaguin_channel_partitioned_consumers{channel=\"events\"} 0\nsaguin_channel_partitioned_consumers{channel=\"jobs\"} 0\nsaguin_channel_partitioned_consumers{channel=\"jobs__dlq\"} 0\nsaguin_channel_partitioned_consumers{channel=\"readings\"} 0\nsaguin_channel_partitioned_consumers{channel=\"state\"} 0\nsaguin_provider_info{provider=\"disk\",type=\"sqlite\"} 1\nsaguin_provider_info{provider=\"mem\",type=\"memory\"} 1\nsaguin_provider_bytes{provider=\"disk\"} 90112\nsaguin_provider_bytes{provider=\"mem\"} 768\nsaguin_provider_max_bytes{provider=\"disk\"} 0\nsaguin_provider_max_bytes{provider=\"mem\"} 0\nsaguin_storage_commits_total{provider=\"disk\",closed_by=\"records\"} 0\nsaguin_storage_commits_total{provider=\"disk\",closed_by=\"interval\"} 0\nsaguin_storage_commits_total{provider=\"disk\",closed_by=\"commit\"} 2\nsaguin_storage_commits_total{provider=\"disk\",closed_by=\"unbatched\"} 0\nsaguin_storage_committed_records_total{provider=\"disk\"} 2\nsaguin_provider_publish_commit_max_records{provider=\"disk\"} 256\nsaguin_bridge_info{bridge=\"head-office\",peer=\"tcp://127.0.0.1:38884\"} 1\nsaguin_bridge_connected{bridge=\"head-office\"} 1\nsaguin_bridge_stopped{bridge=\"head-office\"} 0\nsaguin_bridge_sent_total{bridge=\"head-office\"} 0\nsaguin_bridge_loops_skipped_total{bridge=\"head-office\"} 0\nsaguin_bridge_unsent_total{bridge=\"head-office\",cause=\"peer_refused\"} 0\nsaguin_bridge_unsent_total{bridge=\"head-office\",cause=\"broadcast_refused\"} 0\nsaguin_bridge_unsent_total{bridge=\"head-office\",cause=\"unmappable\"} 0\nsaguin_bridge_unsent_total{bridge=\"head-office\",cause=\"live_queue_full\"} 0\nsaguin_bridge_unsent_total{bridge=\"head-office\",cause=\"superseded\"} 0\nsaguin_bridge_received_total{bridge=\"head-office\"} 2\nsaguin_bridge_unstored_total{bridge=\"head-office\",cause=\"no_rule\"} 0\nsaguin_bridge_unstored_total{bridge=\"head-office\",cause=\"unmappable\"} 0\nsaguin_bridge_unstored_total{bridge=\"head-office\",cause=\"never_accepted\"} 0\nsaguin_bridge_unstored_total{bridge=\"head-office\",cause=\"queue_full\"} 0\nsaguin_bridge_reconnects_total{bridge=\"head-office\"} 0\nsaguin_connections_total 4\nsaguin_bytes_received_total 664\nsaguin_bytes_sent_total 909\nsaguin_publishes_received_total 14\nsaguin_deliveries_sent_total 4\nsaguin_deliveries_refused_total 0\nsaguin_sessions_offline 1\nsaguin_wills_published_total{cause=\"immediate\"} 1\nsaguin_wills_cancelled_total 0\nsaguin_wills_waiting 0\nsaguin_sessions_restored_total 0\nsaguin_sessions_dropped_total{cause=\"storage_full\"} 0\nsaguin_sessions_dropped_total{cause=\"expired_while_stopped\"} 0\nsaguin_sessions_dropped_total{cause=\"subscription_refused\"} 0\nsaguin_retained_messages 0\nsaguin_qos2_held 0\nsaguin_qos2_abandoned_total 0\nsaguin_qos2_max_inflight_per_client 20\nsaguin_session_expiry_shortened_total 0\nsaguin_published_total{channel=\"events\"} 6\nsaguin_published_total{channel=\"jobs\"} 3\nsaguin_published_total{channel=\"jobs__dlq\"} 0\nsaguin_published_total{channel=\"readings\"} 2\nsaguin_published_total{channel=\"state\"} 2\nsaguin_published_total{channel=\"broadcast\"} 1\nsaguin_broadcast_unmatched_total 1\nsaguin_deliveries_dropped_total 0\nsaguin_deliveries_expired_total 0\nsaguin_session_deliveries_dropped_total{cause=\"session_queue_full\"} 0\nsaguin_session_deliveries_dropped_total{cause=\"packet_ids_exhausted\"} 0\nsaguin_session_deliveries_dropped_total{cause=\"no_shared_member\"} 0\nsaguin_session_deliveries_dropped_total{cause=\"storage_full\"} 0\nsaguin_session_deliveries_dropped_total{cause=\"too_large\"} 0\nsaguin_shares_held_total 0\nsaguin_shares_drained_total 0\nsaguin_shares_dropped_total{cause=\"backlog_full\"} 0\nsaguin_shares_dropped_total{cause=\"no_member_left\"} 0\nsaguin_shares_dropped_total{cause=\"expired\"} 0\nsaguin_shares_dropped_total{cause=\"storage_full\"} 0\nsaguin_shares_dropped_total{cause=\"member_ended\"} 0\nsaguin_shares_dropped_total{cause=\"not_authorized\"} 0\nsaguin_session_queue_messages 0\nsaguin_session_queue_bytes 0\nsaguin_publish_refused_total{reason=\"not authorized\"} 1\nsaguin_connections_refused_total{reason=\"bad authentication method\"} 1\nsaguin_channel_retention_removed_total{channel=\"events\"} 0\nsaguin_channel_retention_removed_total{channel=\"jobs\"} 0\nsaguin_channel_retention_removed_total{channel=\"jobs__dlq\"} 0\nsaguin_channel_retention_removed_total{channel=\"readings\"} 0\nsaguin_channel_retention_removed_total{channel=\"state\"} 0\nsaguin_channel_position_lost_total{channel=\"events\"} 0\nsaguin_channel_position_lost_total{channel=\"jobs\"} 0\nsaguin_channel_position_lost_total{channel=\"jobs__dlq\"} 0\nsaguin_channel_position_lost_total{channel=\"readings\"} 0\nsaguin_channel_position_lost_total{channel=\"state\"} 0\nsaguin_latest_superseded_total{channel=\"state\"} 0\nsaguin_queue_depth{channel=\"jobs\"} 3\nsaguin_queue_inflight{channel=\"jobs\"} 1\nsaguin_queue_delivered_total{channel=\"jobs\"} 1\nsaguin_queue_acknowledged_total{channel=\"jobs\"} 0\nsaguin_queue_returned_total{channel=\"jobs\"} 0\nsaguin_queue_redelivered_total{channel=\"jobs\"} 0\nsaguin_queue_dead_lettered_total{channel=\"jobs\"} 0\nsaguin_queue_expired_total{channel=\"jobs\"} 0\nsaguin_queue_retain_ignored_total{channel=\"jobs\"} 0\nsaguin_go_heap_live_bytes 1.0514432e+07\nsaguin_go_heap_goal_bytes 2.1106112e+07\nsaguin_go_gc_cycles_total 3\nsaguin_go_gc_cpu_seconds_total 0.004718377\nsaguin_go_gc_assist_cpu_seconds_total 0.000125117\nsaguin_go_stack_bytes 917504\nsaguin_go_allocated_bytes_total 2.4389552e+07\nsaguin_go_allocated_objects_total 118004\nsaguin_go_gogc_percent 100\nsaguin_go_memory_limit_bytes 9.223372036854776e+18\n```\n\n**One line in it does not read as you might expect.**\n`saguin_provider_bytes{provider=\"mem\"}` is 768 where the channels on it\nsum to 585. The provider holds `state` as well, which `saguin_channel_bytes`\ndoes not measure, and it is `broker.session.storage` here, so the session\nthe durable consumer left is in it too.\n\n### The catalogue\n\nClosed. A name here is a promise; a name not here is not published.\n\n**The broker**\n\n| Metric | Type | |\n|---|---|---|\n| `saguin_build_info{version,broker_id}` | gauge | Always 1. This is where Sagüin's own version is stated, and the reason `$SYS/broker/version` is not to be believed. |\n| `saguin_uptime_seconds` | gauge | |\n| `saguin_connections` | gauge | Connected now. Read from the substrate's own counter, which it maintains whether or not anything asks |\n| `saguin_connections_by_protocol{protocol}` | gauge | Connected now, by the MQTT version each client connected with - `5` or `3.1.1`, and no other label is possible. It is the fleet mix: how much of the estate is still on the older protocol, and whether that number is coming down. **Counted from the client table at each scrape rather than kept by increments**, because a gauge maintained on connect and disconnect is a gauge that drifts. **It sums to `saguin_connections`**: the substrate keeps an inline client of its own in that table - the one it publishes on its own behalf with - and the walk leaves it out, or a broker with one MQTT 5 subscriber would report one connection and two clients speaking MQTT 5. One client is a small and constant error, which is what lets one stand: three hundred against three hundred and one has nothing about it that looks wrong |\n| `saguin_connections_total` | counter | Accepted since start |\n| `saguin_sessions_offline` | gauge | Sessions this broker holds that nothing is connected to. **A session outliving its connection is what a session is for**, so this is a fleet's shape rather than a fault - but it is the question `saguin_connections` cannot answer: three hundred devices with two hundred connected is either a rota or a hundred that have stopped calling, and the connection count alone cannot tell them apart. A number that only climbs is sessions nothing is coming back for, held until their expiry. **A session this broker put back when it started counts here too**, because it is the same thing: a session held with nothing connected to it, whether its client left a minute ago or before the last restart. **Derived rather than read**, because the substrate derives it inside the `$SYS` tree Sagüin silences - the inputs are a map length and a counter, and this scrape already walks that table for the protocol split. |\n| `saguin_wills_published_total{cause}` | counter | Wills this broker published on a client's behalf, by what made each due. `immediate` is a connection that ended without a `DISCONNECT` and no Will Delay Interval; `delayed` is one whose delay ran out; `session_ended` is a session that ended while its Will was still due - it expired, a clean start under its client id ended it, or its client came back to find retention had passed one of its positions (RFC 0003 \"Sessions\"); `start` is a Will this broker found already owed when it started, because the delay passed or the session ended while it was stopped. **Nothing dies because the broker stopped**, so a restart never adds to `immediate` for the clients it disconnects - a fleet's worth of these at a restart would be the broker announcing its own maintenance as a mass outage. |\n| `saguin_wills_cancelled_total` | counter | Wills a client cancelled by resuming its session inside its Will Delay Interval [MQTT-3.1.3-9]. **This is the delay working rather than a fault**: a number climbing here beside a flat `saguin_wills_published_total` is a fleet on a bad link that is correctly not being announced dead, which is the whole reason the interval exists. |\n| `saguin_wills_waiting` | gauge | Wills waiting out their Will Delay Interval now: clients that have gone and whose deaths have not been announced yet. Each is held in `broker.session.storage` with the moment it becomes due, so a restart inside the wait publishes it when it lands rather than losing it - and a number that only climbs is devices leaving faster than their delays run out. |\n| `saguin_sessions_restored_total` | counter | Sessions this broker put back when it started, from what `broker.session.storage` kept while it was stopped. Each is answered `Session Present = 1` when its client comes back and is sent what it was owed. **It moves once, at the start, and then stands still**, so what it reports is what the last start found - read beside `saguin_sessions_dropped_total`, the two say what a restart did to a fleet. A start that restored none, on a broker whose clients ask for persistent sessions, is a session store that did not keep them: a `memory` provider with no `snapshot_dir`, or one that was stopped by a crash rather than a signal. Nothing else says so. |\n| `saguin_sessions_dropped_total{cause}` | counter | Sessions the session store could not keep, by cause. `storage_full` is a client that asked for its session to outlive its connection while the provider `broker.session.storage` names was at its `max_bytes`: it was accepted, and told in its `CONNACK` (Session Expiry Interval 0) that the session ends with the connection. **Nothing else shows a fleet losing its sessions to a full provider**, because the clients connect as usual and only find out when they come back to nothing. `expired_while_stopped` is a session whose expiry passed while the broker was stopped, ended when it started with the messages it was owed. `subscription_refused` is a session ended at a start because it held a subscription this broker's rules refuse - a channel that became a queue while it was stopped, most often, whose records must reach a worker and nobody else. The session goes rather than the subscription, because MQTT has no way to tell a resuming client that one of its subscriptions is gone: it comes back to `Session Present = 0` and is answered for each filter it asks for again. |\n| `saguin_session_queue_messages` | gauge | Deliveries every session holds that its client has not acknowledged, connected or not: on the wire, waiting for room in the client's Receive Maximum, or queued while it is away - for broadcast from the log, every message a session is owed after its cursor (RFC 0003 \"Broadcast\"). **A channel's records waiting to be sent are not in it**: an `append` consumer waits at its position, and a `latest` subscriber's waiting values - at most one per topic its subscriptions reach, a newer one replacing an older - belong to its connection rather than its session, and are bounded by those topics rather than by `limits.session_queue_bytes`. **Summed over sessions and never split by one**, because the only label that would divide it is the client id - a string the client chose. One session's share is bounded by `limits.session_queue_bytes`, and the log names a session when it reaches that. Read from counts each session keeps as deliveries arrive and leave - its in-flight table's, in the same walk as the row above, and its list on the broadcast log - each delivery counted once. |\n| `saguin_session_queue_bytes` | gauge | The memory those deliveries take, as `limits.session_queue_bytes` counts it for one session: payload, topic and properties, and beside them about 600 bytes for each delivery in the client's in-flight table and 300 once for a table holding any, 80 for each message a session is owed from the broadcast log, and both and 192 more for one of those on the wire, 600 once while any is (RFC 0002). |\n| `saguin_retained_messages` | gauge | Retained messages held on broadcast topics, and **Sagüin's own store rather than the substrate's**, whose count is wrong here for the reason given under \"There is no `$SYS` tree\". They accumulate until something publishes an empty payload to the topic or `broker.retained.retention_period` removes them, and nothing else says how many there are. Always reported: every broker keeps the store. |\n| `saguin_qos2_held` | gauge | Exactly-once publishes received and not yet released - messages this broker has taken ownership of at their `PUBREC` and that no client has finished sending. A figure that climbs and stays is publishers not completing exchanges. Each is held in the store of the channel it is for, or the broadcast log's for a broadcast, and the broker answers this from its own count of them, kept as each is held and released, rather than asking every channel's store; their bytes are in their provider's `saguin_provider_bytes`. Always reported, with `saguin_qos2_abandoned_total` and `saguin_qos2_max_inflight_per_client`: every broker offers QoS 2. |\n| `saguin_qos2_abandoned_total` | counter | Exactly-once publishes taken in and never released, because their session ended, they aged past `broker.qos2.expires_after`, or the start found them holding for a session that did not come back or in a channel the configuration does not keep. **Nothing else can show them**: the publisher has had its `PUBREC` and is simply never answered, and one that completes is already counted where its record lands. One series rather than one per reason, because both are the same thing seen by a publisher - an exchange that did not finish. |\n| `saguin_qos2_max_inflight_per_client` | gauge | The configured `broker.qos2.max_inflight_per_client`, published for the reason `saguin_provider_max_bytes` and `saguin_max_session_expiry_seconds` are: a dashboard should be able to read what is held against what was allowed. |\n| `saguin_subscriptions` | gauge | Every subscription the broker holds, not only the ones on channels. The substrate's own counter |\n| `saguin_session_expiry_shortened_total` | counter | CONNECTs that asked to keep their session for longer than `max_session_expiry` and were given the cap. The client is told in its CONNACK; this is what tells the operator, and nothing else does. |\n| `saguin_max_session_expiry_seconds` | gauge | The configured cap, so a dashboard can read what a fleet asks for against what it gets - the same reason `saguin_provider_max_bytes` is published beside the bytes held. |\n\nThe counter is Sagüin's own comparison of the requested Session Expiry\nInterval against the limit, at connect time, on a path that is not the\nhot one.\n\n**What the box moved.** Five counters about the wire rather than about a\nchannel, and the only ones here a throughput is drawn from: everything else\ncounts records Sagüin decided to keep or deliver. Every one of them is read\nfrom a counter the substrate already maintains as it reads and writes\nsockets - the numbers its own `$SYS` tree publishes, which Sagüin silences,\nand silencing stops the publishing and not the counting. Keeping a tally\nbeside one of them would be a second number for one fact.\n\n**They are not split by listener**, and that is a decision rather than a\ngap. The substrate keeps one set of totals for the whole broker, so a\nfigure per door would mean Sagüin counting bytes itself as they cross each\nconnection - the second tally the paragraph above refuses, to answer a\nquestion an operator can usually settle from which addresses are in use.\n\n| Metric | Type | |\n|---|---|---|\n| `saguin_bytes_received_total` | counter | Bytes read from MQTT connections since start. **Everything on the wire**: a payload, the acknowledgement behind it, a keepalive, a connection being set up. That is what sizes a link, and it is not the same question as how much payload a fleet sent. A bridge's traffic to its peer is in none of these four - that link is Sagüin acting as a client through another library, which the substrate's counters never see. |\n| `saguin_bytes_sent_total` | counter | Bytes written to MQTT connections since start, on the same terms. |\n| `saguin_publishes_received_total` | counter | Publishes that arrived, **the ones Sagüin went on to refuse included**: every PUBLISH a client sent, and every record an inbound bridge carried in over its upstream connection. Not a queue's offer to its worker, and not a Will the broker publishes on a client's behalf, which arrived as nothing. Against `saguin_published_total` summed, the pair says how much of what a fleet sends is landing, which neither answers alone. |\n| `saguin_deliveries_sent_total` | counter | Publishes Sagüin sent to subscribers, a queue record handed out a second time included. **This is the read side**, and nothing else in this catalogue counts what leaves a channel that is not a queue. |\n| `saguin_deliveries_refused_total` | counter | **Broadcast** deliveries Sagüin could not send because the subscriber already held `limits.session_queue_bytes` it had not acknowledged, or had run through its packet identifiers. **Every one is also counted in `saguin_session_deliveries_dropped_total`**, under `session_queue_full` or `packet_ids_exhausted`: this is the refusal seen from the delivery, those are the same losses seen from the session, and adding the two counts each twice. A different loss from `saguin_deliveries_dropped_total`, which is that client's outbound queue overflowing: the same complaint one layer apart, and an operator watching only the other sees half of what a slow subscriber costs. **A channel record meeting a full window is neither of them and is not lost** - Sagüin delivers those itself: an `append` record waits at the consumer's position, and a `latest` value waits for that subscriber, replaced by any newer value for its topic (`saguin_latest_superseded_total`), and each is sent when a slot frees. A consumer that leaves first has its position held below what it was not sent, so a resume carries it (RFC 0003). Broadcast promises nothing beyond delivery to whoever is connected at the time, so there is nothing to come back for, which is what makes this half the half worth counting. |\n\n**Publishing**\n\n| Metric | Type | |\n|---|---|---|\n| `saguin_published_total{channel}` | counter | One series per channel, plus one for broadcast. A dead-letter channel is always zero: a record arrives there by the queue's own move and a publish into it is refused, so `saguin_queue_dead_lettered_total` is where its arrivals are counted. |\n| `saguin_connections_refused_total{reason}` | counter | Connections this broker refused, by reason, whatever protocol the client speaks: a CONNECT turned away before a session existed, a connection ended for a refusal, and a socket closed for want of a `max_connections` slot before it sent a CONNECT that could be answered: a websocket before its upgrade, and on any door one arriving with the overflow budget full (RFC 0002). That last is counted as `server busy` whether or not the socket would have sent a CONNECT, and makes no row on `/v1/operations/refused`, since it has no client id; a tcp socket read to be refused is counted when its CONNECT is. An MQTT 5 client is sent a `DISCONNECT` naming the code; one below MQTT 5, whose `PUBACK` carries no reason code, has its connection closed rather than a record acknowledged that Sagüin threw away. Labelled with the same reason vocabulary as the row below, so a rate of one against the other reads directly - of the publishes refused for this reason, how many cost a device its connection. **This is the series behind a fleet in a reconnect loop**, and the one to alert on. `session store full` appears here and nowhere else: a client that armed a Will the provider `broker.session.storage` names had no room for is refused `0x97`, because MQTT has a way to tell a client its session ends with its connection and none to tell it its Will is not held |\n| `saguin_publish_refused_total{reason}` | counter | Labelled with the MQTT specification's name for the code the client was answered with - and never the sentence Sagüin sent beside it. The set is closed but it is not only the specification's: on the publish path `0x97` is the wire answer for six different things, so `publish rate exceeded` and `inflight allowance exceeded` are each split out from `quota exceeded`, and a refusal that carries no code at all is `rejected`. The second of those is a publisher holding as many unfinished exactly-once exchanges as `broker.qos2.max_inflight_per_client` allows it, which is a client to look at where a full channel is storage to look at. This is what answers \"the fleet's data is not arriving\" without reading a log - including the case an operator is most likely to have caused, a rule that refuses the fleet `0x87`, which is counted where the acl_file answers rather than on the publish path it never reaches. The one refusal missing from it is the `$SYS` publish, for the reason given at the end of this document. |\n| `saguin_subscriptions_refused_total{reason}` | counter | Topic filters a `SUBACK` refused, one for each filter, labelled with the MQTT specification's name for the code decided for it. **Counted before a 3.1.1 client's code is written as `0x80`**, the one failure that protocol has, so a 3.1.1 fleet refused by an acl_file reads `not authorized` here and not `unspecified error`. Counted once, where the codes of a `SUBACK` are all known, so every refusal is here whichever part of Sagüin decided it: `topic filter invalid` for a filter that is not one or is deeper than `limits.max_topic_levels`, `not authorized` for a filter the client's roles do not allow, `quota exceeded` for a session store at its bound or a client at `limits.max_subscriptions`, `implementation specific error` for a queue's form, a reserved property or a partition declaration Sagüin refuses, or a session store that could not keep the subscription other than for room, `packet identifier in use`, and `shared subscriptions not supported`. No series until a filter has been refused, as the row above. What answers \"why is this device not getting anything\" when its SUBSCRIBE was answered and nobody read the codes |\n| `saguin_broadcast_unmatched_total` | counter | Publishes that matched no subscription and no channel. **This is what answers a mistyped channel name**, which nothing else can: a publish to `event/x` where the channel is `events` is ordinary broadcast to nobody, and MQTT has no reason code for it. |\n| `saguin_deliveries_dropped_total` | counter | Broadcast deliveries at QoS 0 discarded because a subscriber's outbound queue was full; one at QoS 1 or 2 waits in its session instead. **This is what says a subscriber is being shed**, and nothing else does. It does not count channel deliveries - see below. |\n| `saguin_deliveries_expired_total` | counter | Deliveries discarded because the publisher's Message Expiry Interval ran out before they were sent. One already on the wire is never discarded by expiry, whatever its session, as MQTT has the sender of an exactly-once delivery do (MQTT-4.3.3-7): a resumed session is sent it again under its identifier (MQTT-4.4.0-1), and a session that ends with its connection takes it with it when it ends, which is not counted here. |\n| `saguin_shares_held_total` | counter | Deliveries put on a shared group's list (RFC 0003 \"Broadcast\"): every QoS 1 or 2 delivery a group with a member whose session outlives its connection is owed - one its bound or its provider has no room for as well, counted dropped as it arrives - every one returned to it by a member's ending or dropped there once it had been handed out (`member_ended`, or a return its provider has no room for), and at a start every one put back on a group's list or dropped with a group no session holds. A group over an `append` or `latest` channel is counted exactly as a group over a broadcast topic is: the records it is owed are held as copies in the broadcast log. Everything else is counted as `no_shared_member` below. |\n| `saguin_shares_drained_total` | counter | Deliveries a shared group handed to a member, which is the only way one leaves its list except by being dropped. **Held minus drained minus dropped is what the groups are holding**, exactly; **drained flat while held climbs is a fleet that is not coming back**, and the backlog is spending its provider until one of them does. |\n| `saguin_shares_dropped_total{cause}` | counter | Deliveries a shared group dropped, by a closed set of six causes. `backlog_full`: at `limits.session_queue_bytes` for that group, the oldest given up for a newer one. `no_member_left`: the last member session that could have collected them ended, so nobody was owed them any more - which is the rule that a backlog cannot outlive the sessions it belongs to. `expired`: held longer than `broker.share.expires_after`, dropped by the running sweep or by a start, or past the publisher's own Message Expiry Interval when the group reached it - **a cause of its own, because the label names the knob**, and an expiry counted as a full queue sends an operator to grow a bound that was never reached. `storage_full`: `broker.session.storage` had no room, so the oldest part of the broadcast log went while the group was owed something in it, or a delivery a member's ending returned could not be kept even in the provider's reserve (RFC 0002 \"Every session's state: `broker.session`\"). `member_ended`: a QoS 2 delivery a group handed a member whose session ended before the member answered it with a PUBREC, which MQTT forbids sending to another member (MQTT-4.8.2-5; RFC 0003 \"Broadcast\"). `not_authorized`: the `acl_file` allowed it to none of the members that could take it. **All six series from the start**, at zero until their cause happens, for the reason the five below are. |\n| `saguin_session_deliveries_dropped_total{cause}` | counter | Deliveries a session never received, by a closed set of five causes Sagüin names. `storage_full`: the provider `broker.session.storage` names was full, so the oldest of the broadcast log went and this session was owed it - to make room for the log's next message, or for any other write on a provider the log shares, a channel's record among them (RFC 0002 \"Every session's state: `broker.session`\"). `session_queue_full`: at `limits.session_queue_bytes`, the oldest the session was owed that was not on the wire given up, or a new delivery refused because the session already held that much unacknowledged (RFC 0002). `no_shared_member`: a shared group had no member that could take the message - every member offline, or connected with no room for it. `packet_ids_exhausted`: refused with no packet identifier left. `too_large`: larger than the client's Maximum Packet Size, so discarded rather than sent, as MQTT has the server do [MQTT-3.1.2-25]: a broadcast delivery, whichever way it goes out - live, from the broadcast log, or re-sent to a resumed session - and its subscriber stays connected. A channel's record or a queue's job too large for its client is not counted here: that client is disconnected instead, and a shared group's member it does not fit is passed over (RFC 0003 \"When a record is too large for a subscriber\"). **Five series from the start**, each at zero until its cause happens, because every one is a thing to alert on and a series that appears only once it fires is one a rule cannot be written against in advance. |\n\n**Shedding a slow subscriber is a bound being enforced, not a fault.** A\nclient's outbound queue holds a fixed number of packets, and no more bytes\nthan half its session's `limits.session_queue_bytes`, and the broker\ndiscards past it rather than growing - invariant 13, and the reason one\nsubscriber that has stopped reading cannot become a dead broker. What the\ncounter adds is that it can be seen: it is the only thing that says a\nfleet is being shed.\n\n**It counts broadcast at QoS 0, and that is the whole of it.** A QoS 0\ndelivery reaches the substrate's bounded outbound queue and is discarded\nthere; everything else is held within its session's\n`limits.session_queue_bytes` and given up by that bound's own rules,\ncounted as `session_queue_full` below (RFC 0002). A queue cannot\ncontribute at all - QoS 0 on its form is refused.\n\n**So this counter staying at zero says nothing about a slow consumer.**\nWhat reports one is `session_queue_full`, or\n`disconnected a consumer that stopped reading` at WARN, naming the client\nand the deadline it exceeded - written in `OnDisconnect`, where every route\nends, rather than beside either write. An `append` or `latest` delivery, a\ndead-letter record and every control reply never reach the substrate's\noutbound queue at all: Sagüin writes them to the consumer's socket under\n`limits.write_timeout`, and a consumer that does not take one in time is\n**disconnected** rather than dropped (RFC 0002). So a consumer that stops\nreading those shows up as a disconnection in the log and never here. **A\n`latest` subscriber that reads, but more slowly than its topics change,\nshows up in neither**: each topic's value waits for it and a newer one\nreplaces a waiting one - RFC 0003's promise, the current value rather than\nevery intermediate one - and `saguin_latest_superseded_total` counts what\nwas replaced.\n\n**A queue worker is on that route too, and is disconnected like any other\nclient.** Its offers are written by the substrate, under the same\n`limits.write_timeout`, so a worker that stops reading is hung up on and\nits leases return as any disconnected worker's do. What it does not get is\na lease clock, for the reason invariant 7 gives - a deaf worker and a busy\none look the same from here.\n\n**The two are separate numbers because an operator does a different\nthing about each.** A drop means something is wrong at the other end; an\nexpiry is the publisher's own instruction being carried out, and added\ntogether they would be a figure nobody could act on.\n\n**Neither carries a label, and the useful one is the one that cannot\nexist.** The question is which consumer, and a client picks its own id -\nsee \"Labels, and where the catalogue stops\". The consumer's name is in the\nbroker's log beside the drop instead, bounded the way every client-chosen\nstring is, and written at debug because a consumer thousands of packets in\narrears produces one line per packet.\n\n**Channels.** A channel carries a series here only where the number\nexists. `next_offset` and `floor_offset` are offsets, and a `queue` and a\n`latest` channel have neither - a queue holds work rather than a log, and a\nlatest channel's offsets are sparse, because replacing a value gives it a\nnew one. `records` outlives them both: on an `append` channel it is the\nsubtraction of the two, and on a `latest` channel it is how many topics\nhold a current value. A queue has neither derivation and carries\n`queue_depth` instead. An absent series is deliberate: a zero would read as\nan empty channel rather than as one nobody counts, and a dashboard reading\na wrong number is worse off than one reading nothing.\n`consumer_position_min` follows the same rule for a channel nobody holds a\nposition on, and `records` and `bytes` have their own exceptions below.\n\n| Metric | Type | |\n|---|---|---|\n| `saguin_channel_info{channel,filter,type,provider}` | gauge | Always 1. Carries the topic filter the channel claims, and which storage provider holds it so a dashboard can tell a memory-backed channel from a durable one - see the notes below. |\n| `saguin_channel_records{channel}` | gauge | Records held, by whichever derivation the channel's shape allows. On an `append` channel it is `next − floor`, exact because retention only ever removes a prefix. On a `latest` channel it is how many topics hold a current value - one value per topic, so that count *is* the records - and `next − floor` cannot stand in for it, because a replacement takes a new offset and the subtraction would count writes instead. Both providers answer it from a number they hold rather than by counting rows for the scrape; a stored deletion is a row and counts as one, on both. A `queue` has neither derivation and carries `saguin_queue_depth`. |\n| `saguin_channel_bytes{channel}` | gauge | `append` and `queue` channels. A `latest` channel is absent: one value per topic means a write is a replacement, so measuring it would have to read the value going out first - half again the cost of the write, paid on every publish, to feed a metric. |\n| `saguin_channel_next_offset{channel}` | gauge | |\n| `saguin_channel_floor_offset{channel}` | gauge | The oldest offset still readable |\n| `saguin_channel_retention_removed_total{channel}` | counter | Records retention has deleted |\n| `saguin_channel_consumer_position_min{channel}` | gauge | The lowest stored position. Against `next` it is the worst lag; against `floor` it is the alert below. |\n| `saguin_channel_consumers{channel}` | gauge | How many durable consumers hold a stored position, which is what the lowest one is a reading *from*. Zero carries a series: it is a fact about a channel nobody reads, where an absent one would read as a channel that cannot have consumers. |\n| `saguin_channel_partitioned_consumers{channel}` | gauge | How many subscribers have declared a partition slice on this channel (RFC 0003). It is a denominator and nothing more: it does not say which slices, and it deliberately does not report whether they cover the space - a missing index cannot be told from a member that has not connected yet, so a coverage gauge would be wrong during every rolling restart. The count and the index are not labels, because a client chooses them. |\n| `saguin_channel_position_lost_total{channel}` | counter | Readers whose stored position the retention floor passed, so records they had a claim on were unreadable - invariant 1 firing. **It counts occurrences rather than readers**: the floor passes a lagging reader repeatedly, and a fleet of one straggler can move this number thousands of times. Every kind of reader and every way the loss is discovered, because the question it answers - are records being lost - does not change with which door the loss came through: a consumer overtaken while connected and reading, which MQTT gives no way to tell; one whose stored position had already been passed when it came back, which is told with Session Present = 0; and an outbound bridge rule whose position retention passed. |\n| `saguin_latest_superseded_total{channel}` | counter | On a `latest` channel only: values a subscriber was not sent because a newer value for the same topic took their place while they waited for it - the subscriber had no room, or its delivery was behind (RFC 0003, \"the current value, not every intermediate one\"). **Not a loss**: every subscriber is sent what is current, and settles on it. It is what says a `latest` subscriber reads more slowly than its topics change, which nothing else does - no log line is written for a value replaced, because one per value is how a log stops being read. Against `saguin_published_total{channel}` times the subscribers, it closes the account of what was sent. |\n\n**Queues** - one series per queue channel.\n\n| Metric | Type | |\n|---|---|---|\n| `saguin_queue_depth{channel}` | gauge | Unresolved work |\n| `saguin_queue_inflight{channel}` | gauge | Out with a worker now |\n| `saguin_queue_delivered_total{channel}` | counter | Records handed to a worker, once per hand-over. A record the broker gives back before any worker holds it - no live worker with room, or larger than the chosen worker's Maximum Packet Size - is not counted |\n| `saguin_queue_acknowledged_total{channel}` | counter | |\n| `saguin_queue_returned_total{channel}` | counter | Handed back by a worker |\n| `saguin_queue_redelivered_total{channel}` | counter | Taken back on a visibility timeout |\n| `saguin_queue_dead_lettered_total{channel}` | counter | |\n| `saguin_queue_expired_total{channel}` | counter | Work that aged out unresolved |\n| `saguin_queue_retain_ignored_total{channel}` | counter | Retained publishes this queue took as ordinary work, with the flag dropped. A queue holds work rather than state and grants no subscription a retained message could be delivered to, so there is nothing for the flag to ask for - and nothing on the wire says so, since the publish is acknowledged like any other. It rises when a producer believes it is setting state on a topic a queue claims, which is a question about the operator's own `filter` rather than a broker fault, and this series is the only place it is visible |\n\n**A release lands in exactly one of `returned`, `redelivered` and\n`dead_lettered`, and dead-lettering wins.** A worker's return that spends\nthe last attempt is counted as dead-lettered and not as returned, so\n`delivered = acknowledged + returned + redelivered + dead_lettered` holds\nwith nothing double-counted and nothing missing. `expired` is not one of\nthe three: work that aged out is counted there as well as in whichever of\nthe three released it.\n\n**The identity is over deliveries that have ended**, which is the part to\nread before wiring an alert on it. A delivery in flight has been counted\nin `delivered` and in none of the four, so the two sides differ by\n`saguin_queue_inflight` - which is why the scrape above shows `delivered`\nat 1 beside four zeros and is not a broker miscounting. Subtract the\nin-flight gauge, or compare the two sides on a queue that is idle.\n\n**Storage** - one series per provider.\n\n| Metric | Type | |\n|---|---|---|\n| `saguin_provider_info{provider,type}` | gauge | Always 1; `type` is `memory` or `sqlite` |\n| `saguin_provider_bytes{provider}` | gauge | Everything the provider holds, across every channel on it - including a `latest` channel, which `saguin_channel_bytes` does not measure, so this figure is larger than those summed. On a `sqlite` provider it is the database's own size, which is **not** its disk usage: a write-ahead log waiting to be checkpointed is real bytes on that disk and is not counted here, so do not size a volume from it. It counts freed pages too: retention reuses them rather than returning them, so the figure holds at its peak after a sweep, and that is not retention failing - RFC 0004 says how an operator gives the disk back, on a stopped broker |\n| `saguin_provider_max_bytes{provider}` | gauge | Zero for no bound |\n| `saguin_storage_errors_total{provider}` | counter | Storage calls that failed, and a sqlite provider's periodic fsync of its write-ahead log (`flush_interval`) that failed, which is also logged at ERROR once per streak of failures. No series until a provider has failed, the rule `saguin_publish_refused_total` follows for a code nobody has been answered: a rate over a series that does not exist is nothing, and nothing has gone wrong |\n| `saguin_storage_commits_total{provider,closed_by}` | counter | Transactions that stored publishes, by what closed them: `records` where the transaction filled, `interval` where it did not and waited out `publish_commit_interval`, `commit` where the provider collects without waiting and the transaction held what arrived while the one before it committed, or the one publish that found none committing, `unbatched` on a provider that commits one publish at a time (`none`). A retention sweep and a queue resolution are transactions too and are not counted here. **`sqlite` providers only** - a memory provider has no transaction to collect into |\n| `saguin_storage_committed_records_total{provider}` | counter | Records those transactions carried. `sqlite` only |\n| `saguin_provider_publish_commit_max_records{provider}` | gauge | The provider's `publish_commit_max_records`; 256 where the key is absent and publishes are collected without waiting; zero with `none`, where they are not collected. `sqlite` only |\n\n**Whether collecting publishes is doing anything is the one thing an\noperator cannot see any other way**, which is what those three are for. A\nprovider collecting into shared transactions and one committing singly hold\nidentical records, identical offsets and identical counter rows - the only\ndifference is how long a publisher waited. So:\n\n```\nrate(saguin_storage_committed_records_total[5m])\n / rate(saguin_storage_commits_total[5m])\n```\n\nis the average batch size, and it is read against\n`saguin_provider_publish_commit_max_records` beside it. Where\n`publish_commit_interval` is set, a batch size near one under a ceiling of\nhundreds means every transaction is waiting out the interval and\ncollecting almost nothing, which is **slower than not collecting at all** -\nmeasured in RFC 0002. Without the key, a batch size near one means only\nthat little arrives at once, and it costs nothing: no transaction waits.\nThe same thing shows directly in `closed_by`: where nearly every commit is\nclosed by `interval` rather than by `records`, the count is set above the\ntraffic.\n\n**Bridges** - one series per bridge.\n\n| Metric | Type | |\n|---|---|---|\n| `saguin_bridge_info{bridge,peer}` | gauge | Always 1 |\n| `saguin_bridge_connected{bridge}` | gauge | 1 or 0 |\n| `saguin_bridge_stopped{bridge}` | gauge | 1 when the bridge halted itself rather than losing its link: the peer holds a record larger than this broker's `max_message_size`, so it disconnects rather than deliver it and every reconnection ends the same way on the same record. `saguin_bridge_connected` is 0 beside it - **which is why this gauge exists.** A link that is merely down reads the same on that one, and only this tells an operator whether they are waiting for a reconnection or for a person |\n| `saguin_bridge_received_total{bridge}` | counter | Records that arrived from the peer |\n| `saguin_bridge_unstored_total{bridge,cause}` | counter | Records from the peer the bridge did not store, by a closed set of causes, each a record the peer was told was finished and so lost at this hop, and each logged when it happens - a run of `queue_full` drops once as it begins and once as it ends (RFC 0002 \"Bridges\"). `no_rule`: no inbound rule covers the peer's topic. `unmappable`: a rule covers it and built no topic, or only one in the reserved ` RFC 0005 - Operations | Sagüin documentation space - the bridge's own configuration, the word `saguin_bridge_unsent_total` uses for the same two on the way out. `never_accepted`: Sagüin would refuse the record whatever the channel - more topic or more headers than it allows - so it is dropped rather than retried for ever. `queue_full`: a QoS 0 record dropped because the bridge's queue was full behind a channel that was not taking what it was given. The bridge stores off the connection's read loop, so that a refusing channel cannot keep it from noticing a lost link, and the queue between is bounded: a QoS 1 or 2 record is never dropped, because the peer may hold no more of them unacknowledged than the bridge's Receive Maximum, and a QoS 0 record - which nothing bounds - is dropped past its share of it, at-most-once as its publisher asked. |\n| `saguin_bridge_sent_total{bridge}` | counter | Records this broker forwarded to the peer |\n| `saguin_bridge_loops_skipped_total{bridge}` | counter | Records not forwarded because they arrived over a bridge. **Read it beside `sent_total` or not at all**: a bridge that is connected, sending nothing and skipping everything looks on `sent_total` alone exactly like one with nothing to send, and the two want opposite actions. This one climbing is a topology somebody built - two brokers pointed at each other - where the guard is working and the records are going nowhere |\n| `saguin_bridge_unsent_total{bridge,cause}` | counter | Records an outbound rule did not send the peer, by a closed set of causes split by what an operator does about each. `peer_refused`: the peer refused it for good - its ACL or its limits - logged with the reason code. `broadcast_refused`: the peer refused a broadcast record, or the link was down when it was published - either way it has no store to retry from, so that is the loss. `unmappable`: the rule could build no topic for it, or only one in the reserved ` RFC 0005 - Operations | Sagüin documentation space - the bridge's own configuration, logged. `live_queue_full`: a broadcast record dropped because the rule's queue was full, rather than hold up every publisher behind one slow link. `superseded`: a `latest` value replaced by a newer one for its topic while it waited to cross - **not a loss**, the newer one crosses; it is what says a link ran behind. |\n| `saguin_bridge_reconnects_total{bridge}` | counter | |\n\n**There is no metric for when a memory channel was last written to disk**,\nand its absence is deliberate. Memory durability is exactly the last\nsuccessful snapshot (invariant 14), snapshots are taken at shutdown, and a\ngauge reading \"never, since start\" for the whole life of a healthy process\nwould be read as a fault. `saguin_provider_info{type=\"memory\"}` is the\nhonest form of the same warning: it says which channels are only as\ndurable as the next clean shutdown, which is the fact an operator needs.\n\n### The one alert this exists for\n\nTwo of the channel gauges are the reason the catalogue is worth building:\n\n```\nsaguin_channel_floor_offset > saguin_channel_consumer_position_min\n```\n\n**A consumer's data has already been deleted.** Retention has passed a\nstored position, so when that consumer returns it is refused rather than\nserved the oldest surviving record - invariant 1 holding, which is correct\nand is also a customer's missing afternoon of telemetry. MQTT has no way\nto express it, and a broker that keeps no per-consumer position has no way\nto compute it - Sagüin can because it keeps the two numbers anyway. Its\nwarning shot is `next − consumer_position_min` growing, which is a device\nthat has been offline long enough to be worth looking at before retention\nreaches it.\n\n**The `filter` label is what makes the catalogue usable by a reader rather\nthan only by a dashboard**, because placement is not derivable from the\nname. A tool that scrapes this endpoint learns which channels exist, what\neach is, and which topics each holds. Without it a reader knows a channel\ncalled `water-location` exists and has no way to find out that its records\nare at `iot/water/location/+`, so it cannot subscribe, cannot attribute a\nrecord it receives, and cannot tell an operator where to look. A derived\ndead-letter channel carries its derived filter here for the same reason:\nnobody wrote that one down either.\n\n**It is the filter as written, braces and all**, and that is a choice\nrather than an oversight. A channel keeps one series here - that is what\nmakes this metric a channel list - and a `{a,b}` filter stands for several\nplain ones, so the two cannot both be true of one label. The written form\nis the one that matches the configuration file and the one an operator\nrecognises.\n\n**So a `{a,b}` filter is not something to put on the wire.** A brace is\nconfiguration syntax and not MQTT: subscribing with\n`iot/+/{status,location}/+` is granted, treated as one literal level, and\nmatches nothing for ever - the silent kind. A reader that means to\nsubscribe expands the braces first, as `saguin --route` prints them\nexpanded and as the connectors viewer does before it subscribes.\n\nTwo ways to read the distance, and they answer slightly different questions:\n\n```\non the broker being copied, where the link is a durable consumer:\n saguin_channel_next_offset − saguin_channel_consumer_position_min\n\nacross the pair, comparing one channel on both brokers:\n next_offset{on the source} − next_offset{on the copy}\n```\n\nThe second is only meaningful because a copy stores each record at the\noffset it was given, so the two counters mean the same thing on both\nbrokers - which is the whole reason offsets are preserved. The first is\nsharper where the link is the furthest-behind consumer and needs only one\nbroker scraped, and it says nothing useful where an application consumer is\nfurther behind than the link.\n\nA `latest` channel answers neither, because it keeps no positions and its\noffsets are sparse by nature. What says a copy of one is current is\n`saguin_bridge_connected` and the rate of `saguin_bridge_received_total`.\n\n**The Go runtime**\n\n**The runtime's own figures, not Sagüin's**, read with `runtime/metrics`\nwhen a scrape recomputes the catalogue and at no other time: a few\nmicroseconds a scrape, and nothing while nobody asks. Never with\n`runtime.ReadMemStats`, which stops the world to answer. They are what\nsizes the collector's work: how often it runs, what it costs, and the heap\nit runs against. The live heap counts the 8 MB heap floor (\"The runtime's\nown pauses, and the two variables that move them\"), so a broker holding\nlittle reads about 8 MB there.\n\n| Metric | Type | |\n|---|---|---|\n| `saguin_go_heap_live_bytes` | gauge | Heap the last garbage collection found reachable, the 8 MB heap floor included. From `/gc/heap/live:bytes`. |\n| `saguin_go_heap_goal_bytes` | gauge | Heap size at which the next garbage collection starts. From `/gc/heap/goal:bytes`. |\n| `saguin_go_gc_cycles_total` | counter | Garbage collections completed since start. **Its rate is the figure to watch**: a small heap collecting dozens of times a second costs a fifth of the broker's CPU, which is what the heap floor is for. From `/gc/cycles/total:gc-cycles`. |\n| `saguin_go_gc_cpu_seconds_total` | counter | CPU seconds spent collecting garbage since start, as the Go runtime estimates it. It counts a stop-the-world pause as every processor's time, so it is an upper bound. From `/cpu/classes/gc/total:cpu-seconds`. |\n| `saguin_go_gc_assist_cpu_seconds_total` | counter | The part of that the broker's own goroutines spent helping the collector: work done on a client's or a publisher's path instead of its own. From `/cpu/classes/gc/mark/assist:cpu-seconds`. |\n| `saguin_go_stack_bytes` | gauge | Memory held for goroutine stacks. Two a connection at most - its reader and, once it has been sent something, its writer - so this is most of what an idle connection costs. From `/memory/classes/heap/stacks:bytes`. |\n| `saguin_go_allocated_bytes_total` | counter | Heap bytes allocated since start. From `/gc/heap/allocs:bytes`. |\n| `saguin_go_allocated_objects_total` | counter | Heap objects allocated since start. From `/gc/heap/allocs:objects`. |\n| `saguin_go_gogc_percent` | gauge | The effective `GOGC`: 100 unless set, and -1 when garbage collection is off. From `/gc/gogc:percent`. |\n| `saguin_go_memory_limit_bytes` | gauge | The effective `GOMEMLIMIT`. With none set it reads `math.MaxInt64`, written `9.223372036854776e+18`, which a dashboard can take as no limit. From `/gc/gomemlimit:bytes`. |\n\n### Labels, and where the catalogue stops\n\n**Every label is bounded by something the operator typed.** Channel names,\nprovider names and bridge names come from the configuration file; MQTT\nreason codes are a closed set in the specification. Nothing else qualifies.\n\n**A client id does not qualify**, and this is the line that keeps the\ncatalogue closed for good. A client chooses its own id, so a metric keyed\nby one is a series count chosen by whoever connects - a fleet that\nreconnects with a fresh id per boot writes an unbounded number of series\ninto the operator's monitoring system, which then falls over, some distance\nfrom Sagüin and long after the cause. The same is true of topics and of\nfilters.\n\nSo there is no series per consumer. `saguin_channel_consumers` is\npublished - a count rides the same aggregate - and it is the denominator\nrather than a stand-in: it says a fleet is three hundred, never which of\nthem is behind.\n\n**That distinction is the whole of what a bounded metric can do here.** One\nstraggler in three hundred and a single consumer that has stopped read as\nthe same `consumer_position_min`, and they are a device to replace in the\nmorning against a page tonight. The count separates those two. It does not\nsay *how many* are behind, and no bounded metric can: that needs a\nthreshold nobody can choose on an operator's behalf, or a series per\nconsumer, which is the line above.\n\n### What the `/v1` routes answer with\n\n**This is the one shape Sagüin promises and versions, so it is written\ndown.** `/v1` is in the path because these bodies are a contract: a field\nhere may be added, and a field may not change meaning or quietly disappear -\nthat is what the next number would be for. A dashboard reading `behind`\ntoday has this paragraph to hold the next release to.\n\n`GET /v1/operations/acl?user=\u003cname>` answers the question `saguin --acl`\nanswers at a shell: what may this client do, and which entry decided it.\n`&client_id=\u003cid>` is needed only to resolve a rule written with `%c`.\n\n```json\n{\"user\": \"device-7\", \"acl_file\": \"/etc/saguin/acl.yaml\",\n \"pattern_applied\": \"device-*\", \"patterns_matched\": [\"*\", \"device-*\"],\n \"patterns_in_file\": [\"*\", \"device-*\"], \"client_id_allowed\": true,\n \"grants\": [{\"role\": \"sensor\", \"kind\": \"channel\",\n \"subject\": \"events (iot/+/events/+)\", \"verbs\": [\"write\"], \"denies\": []},\n {\"role\": \"sensor\", \"kind\": \"broker\",\n \"subject\": \"features\", \"verbs\": [], \"denies\": [\"will\"]}],\n \"grants_withheld\": []}\n```\n\n**`denies` is on every grant**, `[]` where the rule takes nothing away, so a\nconsumer ranges over it without checking. Only a `broker: features` rule can\nfill it (RFC 0002 \"Taking a feature away\"), and a denial there holds\nwhatever the client's other roles allow.\n\n**`grants_withheld` is the rules that grant this pair nothing** because\nthe name `%u` or `%c` would put in holds `+`, `#` or `/`, which a topic\nfilter reads as a wildcard or a level (RFC 0002 \"Roles, and users matched\nby pattern\"), as the file writes them; `[]` where none is. Listed rather\nthan left out, so a rule an operator wrote is never silently absent from\nthe answer.\n\n**`client_id_allowed` outranks everything beside it**, and is null where no\n`client_id` was given. An entry's `client_ids` refuses at CONNECT with\n`0x86` - the same code a wrong password gets - before a single rule is\nconsulted, so a body listing what a pair may publish while that pair cannot\nconnect at all answers the wrong question. The reader is somebody whose\ndevice is not working.\n\n**`pattern_applied` is the one that decides**, and the others matched and\ndid nothing: one entry applies, the one spelling the name out most exactly.\nThat is the only way an `acl_file` can take something away without saying\nso, because a shadowed entry is the intended behaviour and cannot be\nrefused at startup - so it is named here, where somebody is already looking\nbecause a device is not doing what its file appears to say.\n\n**A name nothing matches is granted nothing, and that is an answer rather\nthan a `404`.** Users are patterns, so there is no register of real names\nto check a query against - `printer-3` against a file naming `device-*` is\na legitimate question with the answer \"nothing\". The three pattern fields\nare always present for the same reason, null or empty where there is\nnothing to say: a caller reading an empty `grants` can see from\n`patterns_matched` whether it asked about a name the file has no opinion\non, and a shape that changed with the answer would have to be branched on\nbefore it could be read.\n\n**A broker with no `acl_file` answers `{\"acl_file\": null,\n\"everything_allowed\": true}`** rather than an empty grant list, which would\nbe the JSON spelling of \"granted nothing\" about a broker where every\nauthenticated client may do anything.\n\n**The publish limits are not in this body**, although `saguin --acl` prints\nthem. They are a property of the user rather than of a grant, they are in\nthe configuration this listener already serves, and repeating them here\nwould be a second place for them to be wrong.\n\n`GET /v1/operations/config`, `Content-Type: application/json`, is the\nconfiguration this process resolved: the master file and every `!include`\nin one document, every default filled in, and each channel carrying a\n`configured_in` naming the file it was written in.\n\n```json\n{\"broker\": {\"id\": \"edge-1\", \"storage\": {\"default\": \"local\", \"providers\": {\n \"local\": {\"type\": \"sqlite\", \"file_path\": \"/var/lib/saguin/saguin.db\"}}}},\n \"channels\": {\"events\": {\"type\": \"append\", \"filter\": \"events/#\",\n \"storage\": \"local\", \"configured_in\": \"/etc/saguin/saguin.yaml\"}}}\n```\n\n**The resolved values, not the written ones**, which is the whole reason to\nask a running broker rather than read the file: a channel that wrote no\n`filter` has `\u003cname>/#` here, and one that named no storage has the\nbroker-wide default. The configuration file is read once at startup and\nnever again, so this is what the process is running - editing it underneath\ndoes not change this answer.\n\n**What `SIGUSR1` re-reads is not in here.** The log level is a key of this\ndocument and may have moved since it was rendered; the credential files are\nnamed here by path rather than quoted, and their contents are what the\nsignal replaces. `/v1/operations/acl` answers from the file the broker is\nholding, so that route follows a re-read and this one does not.\n\n**`?section=` narrows it, and repeats compose.** `?section=providers`\nanswers `{\"providers\": {…}}`, and `?section=providers§ion=channels`\nanswers with both. The names are `bridges`, `channels`, `limits`,\n`listeners`, `operations`, `providers` and `storage` - an operator's\nvocabulary rather than the document's tree, which is why `providers` is\none of them although it sits under `broker.storage`. **A name that is\nnot one of those is `400`** listing the ones that are, rather than an\nempty object: a section that answered nothing would read as \"this broker\nhas no providers\" to a caller with a typo. A section that is simply\nunused answers with an empty object, which is how those two are told\napart.\n\n**This route is the one place a whole configuration comes out**, so what it\ncannot carry is worth stating twice: no credential is in it, because none\nis in the schema - a password file, an `acl_file` and a bridge's client key\nare all paths. It is behind the same credential as every other `/v1` route,\nand it can be named on its own in a password file's scope field, so a\nscraper can be given `/metrics` without being given this.\n\n**An operator whose entry has no scope field reaches it the day it\nexists**, which is worth saying out loud rather than leaving to be found: a\nuser with no scopes reaches every route, deliberately, so that no password\nfile has to be migrated. On a deployment that has narrowed nobody, adding\nthis route widens what every existing operator can read.\n\n`GET /v1/operations/consumers`, `Content-Type: application/json`:\n\n```json\n{\"channels\": [\n {\"channel\": \"events\", \"next_offset\": 75, \"consumers\": 312, \"returned\": 100,\n \"positions\": [\n {\"reader\": \"mqtt:north-17\", \"offset\": 12, \"behind\": 63,\n \"last_seen\": \"2026-08-28T16:27:27Z\"}\n ]}\n]}\n```\n\n`GET /v1/operations/sessions`, `Content-Type: application/json`:\n\n```json\n{\"sessions\": [\n {\"client_id\": \"north-17\", \"connected\": false, \"user\": \"cohort-north\",\n \"protocol\": 5, \"listener\": \"tcp\", \"remote\": \"10.0.0.7:52104\",\n \"keepalive\": 60, \"expires_after\": 3600, \"clean_start\": false,\n \"subscriptions\": 2,\n \"partitions\": [{\"filter\": \"iot/+/events/+\", \"count\": 3, \"indices\": [0, 2]}]}\n ], \"total\": 312, \"connected\": 310, \"offline\": 2, \"returned\": 100}\n```\n\n**`partitions` is the one place the client-chosen numbers appear**, and it\nis here rather than in a metric because a partition count is declared by\nthe client: *Labels* keeps every series bounded by something the operator\ntyped, and this route is not a series. One entry per filter the\nsession declared a slice on, so the row above reads as \"three slices, and\nthis member holds 0 and 2\" (RFC 0003 \"Client-declared partitioning\").\n\n**A list of objects rather than a map keyed by filter**, so that every key\nin this response is a field name Sagüin chose and every client-chosen\nstring is a value.\n\nThe field is **absent for a session that declared nothing**, which is\nalmost every session, rather than answered as an empty object.\n\n**This is the question the two counts cannot answer.** `saguin_connections`\nbeside `saguin_sessions_offline` says three hundred devices and two hundred\nconnected, which is either a rota or a hundred that have stopped calling.\nThe counts read the same for both; this names them.\n\n**Held sessions come first**, and the list is capped like every other here,\nso what falls off the end is the fleet behaving normally. `expires_after` is\nthe row that matters on one of them: a session nothing is coming back for is\nheld until it runs out, and that is when the number an operator is watching\nwill fall on its own.\n\n**It answers nothing about positions.** `/v1/operations/consumers` already\ndoes, keyed by the same client id, and a second route answering it would be\na second answer to one question kept in step with the first for ever - the\nrule this document draws around records, applied to a number.\n\n**Nothing here acts.** Hanging a client up is a publish to\n`$saguin/sessions/disconnect`, gated by the ACL like every other thing a\nclient may do (RFC 0002 \"Hanging up a client\"), because whether somebody may\nhang up a client is an authorization question and the broker answers those\nin exactly one place. A verb on this listener would need a second answer to\nit.\n\n`GET /v1/operations/users` is who may connect:\n\n```json\n{\"users\": [\"device-7\", \"gateway-1\"], \"anonymous_allowed\": false,\n \"listeners\": {\n \"tcp\": {\"users\": [\"device-7\", \"gateway-1\"], \"anonymous_allowed\": false,\n \"certificate\": \"required\"},\n \"unix\": {\"users\": [\"local-agent\"], \"anonymous_allowed\": true,\n \"certificate\": \"none\"}}}\n```\n\n**The names are what this broker is admitting**, taken from the files it\nloaded at startup rather than read again - so the answer cannot disagree\nwith who actually gets in, which a re-read would the moment somebody edited\na file under a running process.\n\n**The names alone are never the answer, which is why two fields travel with\nthem.** A door admitting anonymous clients lets in names that are in no\nfile at all. A door requiring a client certificate is wrong in both\ndirections at once: none of the names listed can connect, because they hold\npasswords and the handshake wants a certificate, and whoever the authority\nsigned can, with no password and no entry anywhere. A body carrying only\nthe names would read \"closed to all but these two\" about a door standing\nopen to a certificate authority.\n\n**`certificate` is `required`, `accepted` or `none`** - respectively a door\nnothing reaches without one this broker verified, one that checks a\ncertificate where it is offered and takes a password otherwise, and one\nthat examines none. A Unix socket is always `none`: it carries no TLS, and\nits file permissions are what decide who may reach it. What this route\ncannot do is list who holds a certificate, because Sagüin does not know -\nthe authority mints those. Naming the kind of door is the honest answer,\nand it is the one an operator needs to tell \"these names, by password\" from\n\"whoever your CA signed\".\n\n**`listeners` carries every configured door**, each with its own complete\nanswer. A rule listing only the doors that override something reads well\nand is a trap: a door writing `allow_anonymous: true` and no password file\nof its own would be absent under it, and the reader sent to the fields\nabove - which say the opposite about that door. A row per door has no\nfallback to get wrong.\n\n**These are the clients, never the operators.**\n`broker.operations.password_file` names who may read this broker and which\nroutes each of them reaches, and handing that out over the interface it\nguards widens what one leaked credential is worth. No password or hash\nappears on any route, and none can: what is held here is a list of names.\n\n`GET /v1/operations/queues/\u003cchannel>`:\n\n```json\n{\"channel\": \"jobs\", \"unresolved\": 101, \"returned\": 100,\n \"records\": [\n {\"offset\": 4, \"topic\": \"iot/depot/work/w1\", \"attempts\": 2,\n \"state\": \"leased\", \"holder\": \"north-19\",\n \"lease_until\": \"2026-08-28T16:28:02Z\",\n \"first_seen\": \"2026-08-28T16:20:00Z\", \"last_seen\": \"2026-08-28T16:27:44Z\"}\n ]}\n```\n\n`unresolved` is the total; `returned` is how many rows this body\ncarries, and it stops at a hundred. **The two numbers are separate on\npurpose**: a body that returned a hundred rows and said nothing about\nthe total would read as a queue holding a hundred, and the difference\nbetween that and ten thousand is the whole question being asked. `state` is\n`waiting`, `delivering` - handed to a worker whose `PUBACK` has not arrived,\nRFC 0003's `DELIVERING` - or `leased`, and `holder` and `lease_until` are\npresent only when it is `leased`.\n\n**No payload appears in either body**, on any store - the queue route never\nselects one. An operator is asking which records are stuck and what has\nbeen tried, and a body carrying the records themselves would put a fleet's\ndata through an HTTP endpoint that exists to be scraped and logged.\n\nA `reader` is prefixed by what kind of consumer it is, so a durable MQTT\nsession and a bridge are told apart rather than sharing a namespace.\n\n`GET /v1/operations/refused` answers which clients this broker refused, at\nCONNECT or by ending the connection, on any protocol, and why:\n\n```json\n{\"clients\": [\n {\"client_id\": \"tasmota-c4f1\", \"user\": \"fleet\", \"reason\": \"topic name invalid\",\n \"count\": 47, \"last_seen\": \"2026-08-31T09:14:02Z\"}\n ],\n \"returned\": 1, \"tracked\": 1, \"beyond\": 0}\n```\n\n**What these two cover is every connection this broker refused or\nended**, whether or not the client could be told why - deliberately\nwider than the hung-up-on case alone. A device turned away at the door\nfor its protocol version is refused `0x01`, and *is* told, correctly,\nin its own vocabulary; but the operator's question does not change with\nhow well the device was answered. Somebody who has just set\n`min_protocol_version: \"5\"` is turning a fleet away on purpose, and\nthat is exactly the moment they need the list: it is how they find the\ndevices nobody remembered to upgrade.\n\n**A connection refused before its CONNECT named a client** - bytes that\nare not MQTT, a CONNECT that cannot be read, a zero-length client id asked\nto keep a session - is counted like any other, and listed under an empty\n`client_id`, with the user name where the CONNECT got that far.\n\n**Which means Sagüin has to be told, and the telling takes a hook.** A\nCONNECT refused for its version is refused before any Sagüin hook runs,\nso without one the broker never learns which device it was - the only\ntrace is a line in its listener wrapper with no client id.\n`OnConnectRefused` is that hook, and it sits in the MQTT engine rather\nthan in Sagüin's own code because the refusal happens there: by the time\nSagüin could ask, the connection is already gone.\n\n**A counter says how much and never which one**, and that is why this is a\nroute rather than a label. `saguin_connections_refused_total` shows a rate\nthat will not come down; the next question is which of three hundred\ndevices, and a metric must never answer it - a client id is a string a\nclient chose, so a series keyed by one is a series count chosen by whoever\nconnects.\n\n**Worst first, and the cap is the same hundred**, which is what makes the\ncap safe: the rows that fall off the end are the ones nobody was looking\nfor. `returned` is how many rows this body carries and `tracked` how many\nthe broker is holding, for the reason the queue route gives.\n\n`GET /v1/operations/position-lost` answers which readers the retention\nfloor passed, and how much each of them lost:\n\n```json\n{\"readers\": [\n {\"reader\": \"mqtt:truck-114\", \"channel\": \"events\", \"kind\": \"consumer\",\n \"reported\": false, \"count\": 31, \"records_missed\": 18402,\n \"last_position\": 41288, \"last_floor\": 42977,\n \"last_seen\": \"2026-09-23T08:21:45Z\"}\n ],\n \"returned\": 1, \"tracked\": 1, \"beyond\": 0}\n```\n\n**It is the same rule as the route above, applied to the loss invariant 1\nis about.** `saguin_channel_position_lost_total` says a channel passed its\nreaders two thousand times; the operator's next question is which devices\nhave a hole in their history and how big it is, and that is a route rather\nthan a label for the reason given there.\n\n**Rows are keyed by reader and channel**, not by reader alone. A device\nreading three channels can be passed on all three, at different positions\nand for different amounts.\n\n**The reader carries its scheme**, `mqtt:` or `bridge:`, as it does on\n`/v1/operations/consumers` and in the store itself. It is not decoration: a\nclient may call itself `bridge:head-office`, MQTT putting almost no rule on\na client id, and without the scheme that device and the outbound rule of\nthat name would share a row - the counts summed, the kind whichever was\npassed last, and the link's missing records handed to a device.\n\n**`records_missed` accumulates and the positions do not.** How much a\nreader has lost altogether is the question, and it is the order the rows\ncome back in; `last_position` and `last_floor` are the most recent passing\nand explain that row rather than the total. Where `count` is 1 they agree\nexactly - `records_missed` is `last_floor - last_position` - and above 1\nthey do not, because the floor overtakes a lagging reader repeatedly.\n`count` counts those occurrences, never readers.\n\n**`kind` and `reported` are what an operator triages on**, and the three\nkinds are not equally serious:\n\n| `kind` | `reported` | |\n|---|---|---|\n| `consumer` | `false` | Overtaken while connected and reading. MQTT gives no way to tell it, so the operator is the only party who will ever know - the case invariant 1 is written around |\n| `session` | `true` | Its stored position had already been passed when it came back. It is told, with Session Present = 0, so it knows to start fresh rather than reading on over the hole |\n| `bridge` | `false` | An outbound rule's position. The rule resumes at the floor and the peer across the link is told nothing |\n\nA `reported` of `false` is the row to look at first: nothing downstream\nknows those records are missing, so nothing downstream will ask for them\nagain. **It answers for every passing the row aggregates, not the most\nrecent one** - a reader passed while connected and passed again while away\nreads `false`, because the first hole was never mentioned to anybody. A\n`session` row can therefore carry `reported: false`, and that is not a\ncontradiction: the reader was told about the passing it came back to and\nnot about the one before it. `kind` is the most recent passing's.\n\n**The route and the counter reconcile.** Summed over a channel's rows,\n`count` equals `saguin_channel_position_lost_total{channel}` exactly, while\n`beyond` is `0` and within one process lifetime - two instruments on the\nsame events, which is what makes either of them checkable. Once `beyond`\nmoves they part company and stay parted: a displaced row takes its `count`\nwith it, and `beyond` records that a reader did not fit rather than how many\npassings went with it.\n\n**Held in memory, and emptied by a restart**, like the refusal record and\nunlike the stored positions themselves. The counter it stands in front of\nis a process-lifetime number too, and a record that outlived it could never\nbe reconciled against it again. The durable record is the log, which writes\na line at every one of these sites.\n\n**Worst first by `records_missed`, the same hundred returned and the same\nthousand held**, with `beyond` counting what never fitted. When the record\nis full a newcomer displaces the reader that has lost least, and only if it\nhas lost more; when every reader held has lost more than the newcomer, the\nnewcomer is what does not fit. Either way `beyond` moves, because a reader\nshown everything the broker kept and not told that it kept less than it saw\nbelieves it has seen the fleet.\n\n**`beyond` is the one that is easy to leave out.** The record itself is\nbounded - client ids arrive from strangers, and a fleet taking a fresh one\nevery boot would otherwise write an unbounded map into a broker that runs\non one small box (invariant 13). So is each string in it: a client id and a\nuser name are kept to `limits.max_topic_length`, the bound a log line keeps\nthem to, because a CONNECT carries up to 65,535 bytes of each and the\nrecord outlives the connection. When it is full a newcomer displaces a\nclient that has been refused exactly once, never one that has been refused\ntwice, so a device that starts flapping later is still seen and what\nis evicted is always a one-off; when every entry is a repeat offender the\nnewcomer is refused instead. Either way `beyond` counts what did not fit,\nbecause a reader shown everything the broker kept and not told that it kept\nless than it saw believes it has seen the fleet.\n\n### What is deliberately not measured\n\n**Nothing the substrate counts about a path Sagüin does not take.** Two of\nits numbers are wrong here rather than merely irrelevant, and both for the\nsame shape of reason: Sagüin does the work itself and the substrate's\ncounter is only half told.\n\nIts retained count is one - Sagüin holds its own retained store and empties\nthe substrate's as fast as it fills, so that number is about a store nobody\nhas anything in. `saguin_retained_messages` reports Sagüin's instead.\n\n**Its in-flight count is the other, and there is no Sagüin version.** The\nsubstrate raises that number when it sends a QoS 1 publish and lowers it on\nthe acknowledgement. Sagüin sends channel records itself, putting them into\nthe client's in-flight set without raising the count - and the\nacknowledgement lowers it all the same. Measured on a broker that had\ndelivered three records and had them acknowledged: **-3**. A gauge that\ngoes negative is a wrong number, and a wrong number is worse than an absent\none. What a slow broadcast subscriber costs is answered instead by\n`saguin_deliveries_refused_total` and `saguin_deliveries_dropped_total`, and\nwhat a slow `latest` subscriber costs by `saguin_latest_superseded_total` -\ncounters, which only ever rise.\n\n**No histograms, and no latency percentiles.** A timer on the publish path\nis exactly the cost this document exists to refuse: Sagüin's job is MQTT,\nand instrumenting the hot path to observe it is the trade the whole design\ndeclines. Counters and gauges only. If a mean is ever wanted, a count and a\ntotal give one for two atomic adds and no buckets.\n\n**No payloads, ever, anywhere in this catalogue.** A metric is a number,\nand reading records is MQTT's job - the last section says why.\n\n**And nowhere else.** These numbers have one surface. Two of them telling\nthe same story is two things to keep in step and one of them silently\ngoing wrong, which is the reasoning that already refused a third on-disk\nformat and is why there is no `$SYS` tree below.\n\n## Logging, and the process\n\n**Sagüin writes its log to standard output and nowhere else.** It opens no\nfile, rotates nothing, and deletes nothing. Whatever started it is what\ndecides where the output goes and how much of it is kept.\n\n**There is no `log_file`, and that is a decision rather than a gap.**\nInvariant 13: everything that accumulates is bounded, with defined behaviour\nat the bound. A log file accumulates, and it would be the only such thing in\nthe configuration with no bound written beside it - every channel, every\nprovider and every session has one. Bounding it properly means a size, a\nretention count, compression and deletion, which is logrotate reimplemented\ninside a broker, with clock and timezone edges and a broker deleting files it\ndid not open. A dated filename is not an escape: it splits the growth across\nmany files without bounding the total.\n\nAn external `logrotate` rule covers that, so adopting the key would buy the\nsame gap *and* still need the external tool.\n\n**Where the bound actually is, per supervisor:**\n\n| Runs under | Bounded by | Default |\n|---|---|---|\n| systemd → journald | `SystemMaxUse` in `journald.conf` | about 10% of the filesystem, capped at 4G - bounded out of the box |\n| Docker, compose | `max-size` and `max-file` on the logging driver | **unbounded** - must be set |\n| runit → svlogd, daemontools → multilog | the logger's own size and count | bounded |\n| `saguin >> file` | nothing | unbounded, and rotating it under the process loses lines |\n\nThat last row is the cost of declining the key, and it is written here rather\nthan left to be discovered. The answer is the init system, and every\nRaspberry Pi OS ships systemd.\n\n**No logging to a topic.** In a broker with channels a log line about a\ndelivery can cause a delivery.\n\n### The runtime's own pauses, and the two variables that move them\n\n**A broker at full tilt can be paused by its own runtime**, and an operator\nwatching tail latency should know where that comes from. Go collects\ngarbage while the program runs, and a goroutine holding Sagüin's broker-wide\nlock can be made to help with that collection while it holds it. Everything\nwaiting on that lock waits for the collection, so what an operator sees is\nnot a slow consumer but unrelated clients - a subscribe, a disconnect, the\nhealth probe - taking longer than usual while the broker is busy.\n\nMeasured, on the Ryzen 7 260 at about 200,000 deliveries a second to five\nthousand subscribers of one channel: the longest hold of the lock was 96ms,\nof which all but 30ms went with `GOGC=off`. The subscribers ran in the\nbroker's own process, so the collection that held the lock was of their\ngarbage as well as the broker's; a broker serving clients over the network\ncollects only its own.\n\n**`GOGC` and `GOMEMLIMIT` are the two knobs, and they are the Go runtime's\nrather than Sagüin's** - environment variables set where the process is\nstarted, no rebuild and no key in the configuration file. A larger `GOGC`\ncollects less often, which makes those pauses rarer; `GOMEMLIMIT` is what\nstops that costing unbounded memory, by giving the runtime a ceiling to\ncollect against. `GOGC=400 GOMEMLIMIT=6GiB` is the shape of it, with the\nceiling chosen from the machine and what the deployment holds - RFC 0004's\nfigures for the stores, plus what the sessions and in-flight windows come to.\nWhat the collector is doing is on `/metrics`: `saguin_go_gc_cycles_total`,\nits CPU and the heap it runs against (\"The Go runtime\" in the catalogue).\nSagüin also holds a fixed 8 MB minimum-heap allowance, so a broker with few\nconnections does not collect many times a second: it costs about 1 MB of\nresident memory at rest and up to 10 MB under load, it counts toward\n`GOMEMLIMIT` like any heap, and with `GOGC` raised it matters even less.\n\n**It tunes a probability, not a guarantee**, and it moves memory rather than\nremoving a bound: with `GOGC` raised, the bound on what Sagüin's process\nholds is the one the operator sets in `GOMEMLIMIT`, and a ceiling not set is\na ceiling not there. Sagüin's own bounds - a channel's `max_bytes`, a\nsession's queue, an in-flight window - are unaffected either way, and they\nremain where the broker's own limits are written.\n\n### `broker.pid_file`\n\n**A pid file is offered where a log file is not, and the difference is\nthat it cannot grow**: it holds one number and is replaced. RFC 0002\n\"Where the process writes its id\" has the key and its rules.\n\n### `SIGUSR1` re-reads the log level\n\n**And the credential files, the `ws` listener's `same_origin` and\n`allowed_origins`, and every TLS listener's certificate and client CA,\nwhich are the only other things this signal touches** - each because a\nrestart on a broker holding durable sessions costs every connection and\nevery in-flight record (RFC 0002 \"How much the broker says\" and\n\"Withdrawing a device's access\").\n\n**The signal says what it did not do.** It logs the level it applied, the\n`same_origin` it applied and how many sites `allowed_origins` now lists,\nthe credential files it re-read and what they now hold, each certificate\nit re-read and when that one expires, and states that addresses and\nstorage are startup-only - so an operator who edited `max_message_size`\nand sent the signal sees that it was not applied rather than believing it\nwas.\n\n**An invalid or unreadable file changes nothing and does not stop the\nbroker.** The operator may be mid-edit, and a signal that took a running\nbroker down over a half-saved file would be worse than no signal.\n\n**`SIGHUP` is caught and does nothing**, at `warn`, saying there is no\nconfiguration reload and what `SIGUSR1` re-reads instead - RFC 0002 \"How\nmuch the broker says\" has why an unhandled `HUP` would be worse than a\nno-op and why a silent no-op would be worse than a line.\n\n**There is no `SAGUIN_LOG_LEVEL`.** A live signal makes an environment\nvariable a third way to set one thing, and a precedence table nobody needs.\n\n## There is no `$SYS` tree\n\n**Sagüin publishes nothing under `$SYS`.** Broker statistics come from\n`/metrics` and from nowhere else. There is no configuration key for this,\nbecause there is no feature to switch on.\n\nThree reasons, in order of weight.\n\n**It cannot be put behind a credential.** Everything else Sagüin tells an\noperator sits on an authenticated listener. `$SYS` is read over MQTT by\nany client that can connect, and MQTT's authorization is topic access - so\ngranting it means granting it to devices, and refusing it means a\nsubscription refusal a fleet client meets at runtime. There is no third\nsetting. `/metrics` asks who is asking; a topic cannot.\n\n**It costs on a timer rather than on demand.** Such a tree is refreshed on\nan interval from startup to shutdown, reading process memory statistics\nand republishing a retained message per counter, whether or not one client\nhas ever subscribed. Everything else in this document is paid for when\nsomebody asks for it, and on a box whose job is MQTT that difference is\nthe point.\n\n**Two surfaces for one set of numbers is two things to keep in step**, and\none of them going quietly wrong. The catalogue above is the interface.\n\n**The substrate underneath Sagüin publishes such a tree of its own, and\nSagüin silences it.** Leaving it would be worse than either answer,\nbecause two of its counters are wrong about Sagüin rather than merely\nirrelevant: its `version` names the substrate, and its `retained` counts a\nstore Sagüin empties as fast as it fills, so it reports the tree's own\nentries and never Sagüin's. A dashboard reading a wrong number is worse\noff than one reading nothing, and the first thing anybody points at a new\nbroker is `$SYS/broker/version`.\n\n**A subscription to any topic under `$SYS/` is refused with `0x87 Not\nauthorized`**, rather than granted and then never delivered to: a\nsubscription that will never carry anything is a promise Sagüin cannot\nkeep. It applies to a wildcard filter and to a literal topic name alike.\n\n**A publish to `$SYS/` is refused with the same code and is the one\nrefusal `saguin_publish_refused_total` does not count.** The substrate\nanswers it before any hook runs, so the packet never reaches Sagüin at all -\nwhich is right, and means there is no log line either: this refusal is\ncounted nowhere and written nowhere. It is the whole of the exception:\nevery other refusal in this document passes through the publish path and\nis counted there.\n\nNothing under `$SYS` is part of Sagüin's interface, and nothing there is\nto be believed.\n\n## Statistics over HTTP, records over MQTT\n\n**Nothing on this listener returns the contents of a message.** An\noperator who may read a channel reads it with an ordinary MQTT client: an\n`append` channel, a `latest` channel, a dead-letter channel and broadcast\ntraffic are ordinary MQTT topics and need nothing from this document. A\nlive queue is the one that answers differently, and it answers the same\nway to everybody - it is consumed only through `$saguin/queue/` followed\nby the queue's name, at QoS 1, and every other subscription form is\nrefused, so a plain `jobs/#` gets a subscription failure rather than\nsomebody's work.\n\nThe reason it is drawn here rather than a route being added: whether a\ngiven person may read a given channel is an authorization question, and\nthe broker answers it in exactly one place for every client\n(invariant 10). A second route reading records would need a second answer\nto the same question, kept in step with the first for ever.\n"}