Yggdrasil reachability monitor with whitelist detection, SQLite storage and a zero-dependency dashboard.
  • Python 46.1%
  • HTML 41.4%
  • Shell 12.5%
Find a file
Gluek e2979bac91 Race address candidates within one deadline and report per-address reasons
Candidates used to be tried one after another with the full timeout each, so a
blackholed IPv6 address plus a second slow address overran the 12 s phase budget
and the probe was reported only as "did not finish (DNS hang?)".

- try the first two candidates happy-eyeballs style (0.3 s apart, or right away
  after a fast failure), first success wins, one shared deadline per probe
- failures list every attempt: "v6 2a00:..: connect timeout; v4 1.2.3.4: tls timeout"
- a probe that still overruns says where it was stuck (dns / connect / tls handshake)
- dashboard: peers table shows the short reason from the new detail format

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Y9etgKBLZJf9qJD86wGmDU
2026-09-29 00:25:28 +00:00
systemd Serve the dashboard over Yggdrasil with a source-address allowlist 2026-09-29 00:14:30 +00:00
.gitignore ygg-mon: Yggdrasil connectivity monitor for whitelist conditions 2026-09-28 23:30:33 +00:00
config.example.json Prefer IPv4 on hosts without IPv6; replace gosuslugi.ru canary 2026-09-29 00:19:05 +00:00
dashboard.html Race address candidates within one deadline and report per-address reasons 2026-09-29 00:25:28 +00:00
install.sh install.sh: keep existing web settings on re-runs 2026-09-29 00:19:50 +00:00
LICENSE Add MIT license 2026-09-28 23:40:59 +00:00
nftables-ygg.nft Serve the dashboard over Yggdrasil with a source-address allowlist 2026-09-29 00:14:30 +00:00
README.md Prefer IPv4 on hosts without IPv6; replace gosuslugi.ru canary 2026-09-29 00:19:05 +00:00
ygg_mon.py Race address candidates within one deadline and report per-address reasons 2026-09-29 00:25:28 +00:00

ygg-mon: connectivity over Yggdrasil, with and without whitelists

A kit for one or more "probes". It installs Yggdrasil with your peers, checks every minute whether packets reach ygg.chat.gluek.info (RU) and ygg.chatmail.uk (UK), and at the same time works out which mode the network is in: open internet, whitelist, or no network. Everything goes into SQLite, and the dashboard shows availability separately for each mode.

Dependencies: Python ≥ 3.9 (stdlib only), ping, yggdrasilctl. The dashboard uses no external CDNs, so it loads under whitelists too.

Data collection scheme

                    ┌──────────── probe (VPS / mini-PC on LTE / laptop) ────────────────┐
                    │  systemd timer (1/min) → ygg_mon.py collect → SQLite (WAL)        │
                    │                                   ↑                               │
                    │           ygg_mon.py serve ───────┘ → http://127.0.0.1:8787/      │
                    └──────┬──────────────┬──────────────────┬───────────────┬─────────┘
                           │              │                  │               │
   ① canaries (TLS)        │   ② underlay (TCP/TLS)  ③ ygg sessions      ④ overlay (ping6, TCP)
   ya.ru, vk.com,…  ◄──────┤   peers' host:port ◄─── yggdrasilctl ────► nodes' 200::/7 addrs
   google, 1.1.1.1,… ◄─────┘   (bypassing Ygg)       getPeers            ygg.chat.gluek.info
         │                            │                  │               ygg.chatmail.uk
         ▼                            ▼                  ▼                     │
   mode: open /               "is the IP let      "does the ygg        "do packets reach
   whitelist / offline        through?" SYN       protocol come up?"    the node?" + ⑤ DNS AAAA
                              dropped or RST

Each run is one row in runs (time, mode, daemon state, how many peers are up) and about 20 rows in checks, one per check:

Layer target What exactly ok = 1 means
① canary canary:ru:ya.ru:443, canary:foreign:* … TLS handshake with certificate validation the site opens
② underlay underlay:tls://chatmail.uk:4313 TCP connect (+TLS for tls://) straight to the peer the peer's IP:port is let through
③ peer peer:quic://chatmail.uk:4315 yggdrasilctl getPeers → up, latency, last_error the Yggdrasil session is up
⑤ dns dns:ygg.chatmail.uk AAAA in 200::/7 the name resolves
④ overlay node:uk:icmp, node:uk:tcp443 ping6 ×3 and TCP connect to the ygg address a reply came back (an RST from a closed port counts)
summary node:ru, node:uk any overlay check succeeded packets reach the node

Network mode is decided from the canaries on every run: any foreign site reachable → open; only Russian ones → whitelist; nothing → offline. You can adjust the lists in the config (for example, add sites you know are on your carrier's whitelist).

Failure reasons are stored in detail, split by stage, and that is the diagnostics: connect timeout (SYN dropped, typical for whitelists), tls reset (RST) / tls timeout (DPI cuts the handshake), dns: …, refused. Successful checks store no detail.

Reading the combinations

What you see Likely cause
whitelist, underlay → UK connect timeout, UK peer down, but UK node ok traffic to UK transits the RU node: the chat.gluek.info peer is alive and the RU↔UK link works
underlay ok, but the peer on that port is down the IP is let through, but the protocol itself is cut (DPI on the TLS/QUIC fingerprint)
underlay connect timeout on every peer the peers' IPs aren't whitelisted; you need a peer with a whitelisted IP
peer up, node fail a problem in the overlay or on the node, not the last mile
DNS fail, node ok DNS died under the whitelist; the check runs on the cached address
QUIC up while TCP/TLS are dead (or the reverse) protocol-based filtering; shows which transport survives better

Where to put a probe

Whitelists are switched on mainly on mobile networks, so the main probe should sit on a mobile link: a mini-PC / Raspberry Pi with an LTE modem or USB dongle, or an old laptop on a phone hotspot. A VPS in a Russian datacenter (ru.gluek.info) is useful as a control probe: there are usually no whitelists there, but you can see TSPU filtering between the datacenter and abroad. Probes can be merged into one database (see "Multiple probes").

Installation

git clone https://git.gluek.info/gluek/ygg-mon.git && cd ygg-mon
sudo ./install.sh --probe ru-mts-lte            # probe name as shown on the dashboard
sudo ./install.sh --probe ru-mts-lte --firewall # + block inbound to our ygg address

Clone while the probe still has open internet: if git.gluek.info is hosted abroad, it won't be reachable from a probe that is already under a whitelist.

To update later:

cd ygg-mon && git pull && sudo ./install.sh --probe ru-mts-lte

The installer keeps your existing /etc/ygg-mon/config.json and writes the new default next to it as config.json.new.

What install.sh does (Debian 11+/Ubuntu 22.04+):

  1. installs python3, iputils-ping, curl, gnupg;
  2. installs Yggdrasil: the official apt repository (key checked against fingerprint 1C5162E1…6011C5EA) → the distro package (Ubuntu 24.04 ships 0.5.5) → a .deb from GitHub. If none of these is reachable from Russia, download the .deb in advance and pass --deb ./yggdrasil-….deb. An existing Yggdrasil is not reinstalled, and --no-ygg leaves it alone entirely;
  3. adds your peers to /etc/yggdrasil/yggdrasil.conf (keys and everything else are untouched, a backup is kept next to it, and the new config is validated with yggdrasil -address before it is swapped in);
  4. puts ygg_mon.py + dashboard.html in /opt/ygg-mon, the config in /etc/ygg-mon/config.json, and the database in /var/lib/ygg-mon/;
  5. enables ygg-mon-collect.timer (every minute) and ygg-mon-web.service;
  6. with --firewall, adds a separate nft table ygg_guard: new inbound connections to 200::/7 are dropped, ICMPv6 is allowed. Your ruleset/ufw/docker are not touched. Don't use this on the relays themselves: they need inbound over Yggdrasil.

A probe on Yggdrasil is reachable by the whole network: anything listening on :: is exposed on its 200::/7 address. That's why the dashboard listens on 127.0.0.1 by default, and on a probe that needs nothing else you should enable --firewall.

Default peers:

tls://chat.gluek.info:7726
tls://chatmail.uk:4313
tcp://chatmail.uk:4314
quic://chatmail.uk:4315

A different set: --peers "tls://a:1,quic://b:2".

Usage

ygg-mon status                                  # latest run in the terminal
ygg-mon collect                                 # run the checks by hand, with output
ygg-mon annotate --from 40m "MTS: whitelist on" # mark the timeline
ygg-mon annotate --from "2026-09-28 18:00" --to "2026-09-28 19:30" "router reboot"
ygg-mon report -o ygg.html                      # static snapshot (6h/24h/7d/30d inside)
ygg-mon json --window 7d                        # summary JSON for scripts
journalctl -u ygg-mon-collect -n 50

Dashboard: ssh -L 8787:127.0.0.1:8787 probe-host → http://127.0.0.1:8787/, or directly over Yggdrasil (below). It refreshes every minute.

Dashboard over Yggdrasil

The dashboard can also be served on the probe's own ygg address, so you can open it from any of your devices on Yggdrasil without an SSH tunnel, including while the probe is under a whitelist, as long as ygg itself gets through.

sudo ./install.sh --probe ru-vps --firewall --web-ygg --allow 200:aaaa:bbbb:cccc::1
# then open http://[<probe ygg address>]:8787/

There is no login. Access is limited to the addresses in --allow (comma-separated, single addresses or prefixes such as 300:aaaa:bbbb:cccc::/64 for a whole ygg subnet). This works because a Yggdrasil source address is derived from the sender's public key and can't be spoofed. The allowlist is enforced twice:

  • by ygg-mon itself: any other client gets 403 (loopback is always allowed, so ssh -L keeps working);
  • with --firewall, by nftables: only the allowed addresses can reach the port at all.

Other details:

  • the service listens on 127.0.0.1 and on our own ygg address only, never on ::, so it isn't exposed on a public IPv6 address by accident;
  • at boot it waits up to 2 minutes for yggdrasil to report our address;
  • summaries are cached for cache_seconds (30) and computed one at a time, so repeated requests don't load a small VPS;
  • re-running install.sh with new --allow values updates only the web section of an existing config and regenerates the nft rule.

Find a device's ygg address with yggdrasilctl getSelf.

What's on the dashboard:

  • current mode and the share of time spent in each mode;
  • node cards: whether packets get through, RTT, loss, and availability in "open" and "whitelist" mode separately;
  • timeline: a mode strip + a strip per check (green / partial / none);
  • ping6 RTT to both nodes, with whitelist periods shaded;
  • "availability by network mode" table: every layer × every mode, the main answer;
  • peers now (uptime, latency, traffic, last error) and whitelist episodes with node availability inside each one.

Data and SQL

/var/lib/ygg-mon/ygg-mon.db, tables runs, targets, checks, annotations, kv, plus the v_checks view that joins them. About 1 MB per day, kept for retention_days (60).

-- node availability by mode over the last week
SELECT target, mode, COUNT(*) n, ROUND(100.0*SUM(ok)/COUNT(*),1) pct, ROUND(AVG(rtt_ms),1) rtt
FROM v_checks WHERE layer='node' AND ts > strftime('%s','now','-7 days')
GROUP BY target, mode;

-- why peers failed to come up under whitelists
SELECT target, detail, COUNT(*) FROM v_checks
WHERE layer IN ('underlay','peer') AND mode='whitelist' AND ok=0
GROUP BY target, detail ORDER BY 3 DESC;

Multiple probes

Each probe has its own database and name (probe). To see them in one dashboard, collect the databases on one machine:

scp probe-lte:/var/lib/ygg-mon/ygg-mon.db /tmp/lte.db
ygg-mon import /tmp/lte.db     # idempotent, duplicates are skipped

A probe selector then appears on the dashboard.

Config

/etc/ygg-mon/config.json, see config.example.json. The main keys:

  • nodes[].host: the node's ygg name; the last good address is cached, and you can set addr by hand as a fallback;
  • nodes[].tcp_ports: ports for the TCP check through the overlay (an RST also counts as ok);
  • canaries.ru / canaries.foreign: how the mode is detected. Prefer sites that answer from any network: some (gosuslugi.ru, for example) filter traffic and may not answer from every network;
  • ip_family: auto (default) tries IPv4 only when the host has no global IPv6 route, otherwise alternates IPv6/IPv4; 4 or 6 forces one family for canaries and underlay;
  • ygg_endpoint: if yggdrasilctl can't find the socket by itself (unix:///var/run/yggdrasil/yggdrasil.sock);
  • web.listen: an address, "ygg" (our own ygg address) or a list, e.g. ["127.0.0.1", "ygg"]; web.allow: ygg addresses/prefixes allowed to open the dashboard (empty = anyone who can reach the listen address).

Limitations

  • A QUIC peer can't be checked directly with the stdlib; only the ygg session state covers it.
  • The checks cover handshakes and ping, not throughput: the TLS "freeze" after ~16 KB that TSPU applies isn't visible this way. It shows up indirectly as a peer that comes up and drops right away (last_error, uptime).
  • Mode detection from canaries is a heuristic. If your carrier's whitelist differs, adjust canaries.ru.
  • macOS: the collector works (ping6), but there's no installer; run ygg_mon.py collect from launchd/cron and pass --config.

License

MIT, see LICENSE.