- Python 46.1%
- HTML 41.4%
- Shell 12.5%
Candidates used to be tried one after another with the full timeout each, so a blackholed IPv6 address plus a second slow address overran the 12 s phase budget and the probe was reported only as "did not finish (DNS hang?)". - try the first two candidates happy-eyeballs style (0.3 s apart, or right away after a fast failure), first success wins, one shared deadline per probe - failures list every attempt: "v6 2a00:..: connect timeout; v4 1.2.3.4: tls timeout" - a probe that still overruns says where it was stuck (dns / connect / tls handshake) - dashboard: peers table shows the short reason from the new detail format Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Y9etgKBLZJf9qJD86wGmDU |
||
|---|---|---|
| systemd | ||
| .gitignore | ||
| config.example.json | ||
| dashboard.html | ||
| install.sh | ||
| LICENSE | ||
| nftables-ygg.nft | ||
| README.md | ||
| ygg_mon.py | ||
ygg-mon: connectivity over Yggdrasil, with and without whitelists
A kit for one or more "probes". It installs Yggdrasil with your peers, checks every
minute whether packets reach ygg.chat.gluek.info (RU) and ygg.chatmail.uk (UK),
and at the same time works out which mode the network is in: open internet,
whitelist, or no network. Everything goes into SQLite, and the dashboard shows
availability separately for each mode.
Dependencies: Python ≥ 3.9 (stdlib only), ping, yggdrasilctl. The dashboard uses no
external CDNs, so it loads under whitelists too.
Data collection scheme
┌──────────── probe (VPS / mini-PC on LTE / laptop) ────────────────┐
│ systemd timer (1/min) → ygg_mon.py collect → SQLite (WAL) │
│ ↑ │
│ ygg_mon.py serve ───────┘ → http://127.0.0.1:8787/ │
└──────┬──────────────┬──────────────────┬───────────────┬─────────┘
│ │ │ │
① canaries (TLS) │ ② underlay (TCP/TLS) ③ ygg sessions ④ overlay (ping6, TCP)
ya.ru, vk.com,… ◄──────┤ peers' host:port ◄─── yggdrasilctl ────► nodes' 200::/7 addrs
google, 1.1.1.1,… ◄─────┘ (bypassing Ygg) getPeers ygg.chat.gluek.info
│ │ │ ygg.chatmail.uk
▼ ▼ ▼ │
mode: open / "is the IP let "does the ygg "do packets reach
whitelist / offline through?" SYN protocol come up?" the node?" + ⑤ DNS AAAA
dropped or RST
Each run is one row in runs (time, mode, daemon state, how many peers are up) and
about 20 rows in checks, one per check:
| Layer | target |
What exactly | ok = 1 means |
|---|---|---|---|
① canary |
canary:ru:ya.ru:443, canary:foreign:* … |
TLS handshake with certificate validation | the site opens |
② underlay |
underlay:tls://chatmail.uk:4313 |
TCP connect (+TLS for tls://) straight to the peer | the peer's IP:port is let through |
③ peer |
peer:quic://chatmail.uk:4315 |
yggdrasilctl getPeers → up, latency, last_error |
the Yggdrasil session is up |
⑤ dns |
dns:ygg.chatmail.uk |
AAAA in 200::/7 | the name resolves |
④ overlay |
node:uk:icmp, node:uk:tcp443 |
ping6 ×3 and TCP connect to the ygg address | a reply came back (an RST from a closed port counts) |
| summary | node:ru, node:uk |
any overlay check succeeded | packets reach the node |
Network mode is decided from the canaries on every run: any foreign site reachable
→ open; only Russian ones → whitelist; nothing → offline. You can adjust the lists
in the config (for example, add sites you know are on your carrier's whitelist).
Failure reasons are stored in detail, split by stage, and that is the diagnostics:
connect timeout (SYN dropped, typical for whitelists), tls reset (RST) / tls timeout
(DPI cuts the handshake), dns: …, refused. Successful checks store no detail.
Reading the combinations
| What you see | Likely cause |
|---|---|
whitelist, underlay → UK connect timeout, UK peer down, but UK node ok |
traffic to UK transits the RU node: the chat.gluek.info peer is alive and the RU↔UK link works |
| underlay ok, but the peer on that port is down | the IP is let through, but the protocol itself is cut (DPI on the TLS/QUIC fingerprint) |
underlay connect timeout on every peer |
the peers' IPs aren't whitelisted; you need a peer with a whitelisted IP |
| peer up, node fail | a problem in the overlay or on the node, not the last mile |
| DNS fail, node ok | DNS died under the whitelist; the check runs on the cached address |
| QUIC up while TCP/TLS are dead (or the reverse) | protocol-based filtering; shows which transport survives better |
Where to put a probe
Whitelists are switched on mainly on mobile networks, so the main probe should
sit on a mobile link: a mini-PC / Raspberry Pi with an LTE modem or USB dongle, or an
old laptop on a phone hotspot. A VPS in a Russian datacenter (ru.gluek.info) is
useful as a control probe: there are usually no whitelists there, but you can see
TSPU filtering between the datacenter and abroad. Probes can be merged into one
database (see "Multiple probes").
Installation
git clone https://git.gluek.info/gluek/ygg-mon.git && cd ygg-mon
sudo ./install.sh --probe ru-mts-lte # probe name as shown on the dashboard
sudo ./install.sh --probe ru-mts-lte --firewall # + block inbound to our ygg address
Clone while the probe still has open internet: if git.gluek.info is hosted abroad, it won't be reachable from a probe that is already under a whitelist.
To update later:
cd ygg-mon && git pull && sudo ./install.sh --probe ru-mts-lte
The installer keeps your existing /etc/ygg-mon/config.json and writes the new
default next to it as config.json.new.
What install.sh does (Debian 11+/Ubuntu 22.04+):
- installs
python3,iputils-ping,curl,gnupg; - installs Yggdrasil: the official apt repository (key checked against fingerprint
1C5162E1…6011C5EA) → the distro package (Ubuntu 24.04 ships 0.5.5) → a.debfrom GitHub. If none of these is reachable from Russia, download the.debin advance and pass--deb ./yggdrasil-….deb. An existing Yggdrasil is not reinstalled, and--no-yggleaves it alone entirely; - adds your peers to
/etc/yggdrasil/yggdrasil.conf(keys and everything else are untouched, a backup is kept next to it, and the new config is validated withyggdrasil -addressbefore it is swapped in); - puts
ygg_mon.py+dashboard.htmlin/opt/ygg-mon, the config in/etc/ygg-mon/config.json, and the database in/var/lib/ygg-mon/; - enables
ygg-mon-collect.timer(every minute) andygg-mon-web.service; - with
--firewall, adds a separate nft tableygg_guard: new inbound connections to 200::/7 are dropped, ICMPv6 is allowed. Your ruleset/ufw/docker are not touched. Don't use this on the relays themselves: they need inbound over Yggdrasil.
A probe on Yggdrasil is reachable by the whole network: anything listening on
::is exposed on its 200::/7 address. That's why the dashboard listens on127.0.0.1by default, and on a probe that needs nothing else you should enable--firewall.
Default peers:
tls://chat.gluek.info:7726
tls://chatmail.uk:4313
tcp://chatmail.uk:4314
quic://chatmail.uk:4315
A different set: --peers "tls://a:1,quic://b:2".
Usage
ygg-mon status # latest run in the terminal
ygg-mon collect # run the checks by hand, with output
ygg-mon annotate --from 40m "MTS: whitelist on" # mark the timeline
ygg-mon annotate --from "2026-09-28 18:00" --to "2026-09-28 19:30" "router reboot"
ygg-mon report -o ygg.html # static snapshot (6h/24h/7d/30d inside)
ygg-mon json --window 7d # summary JSON for scripts
journalctl -u ygg-mon-collect -n 50
Dashboard: ssh -L 8787:127.0.0.1:8787 probe-host → http://127.0.0.1:8787/,
or directly over Yggdrasil (below). It refreshes every minute.
Dashboard over Yggdrasil
The dashboard can also be served on the probe's own ygg address, so you can open it from any of your devices on Yggdrasil without an SSH tunnel, including while the probe is under a whitelist, as long as ygg itself gets through.
sudo ./install.sh --probe ru-vps --firewall --web-ygg --allow 200:aaaa:bbbb:cccc::1
# then open http://[<probe ygg address>]:8787/
There is no login. Access is limited to the addresses in --allow (comma-separated,
single addresses or prefixes such as 300:aaaa:bbbb:cccc::/64 for a whole ygg subnet).
This works because a Yggdrasil source address is derived from the sender's public
key and can't be spoofed. The allowlist is enforced twice:
- by ygg-mon itself: any other client gets
403(loopback is always allowed, sossh -Lkeeps working); - with
--firewall, by nftables: only the allowed addresses can reach the port at all.
Other details:
- the service listens on
127.0.0.1and on our own ygg address only, never on::, so it isn't exposed on a public IPv6 address by accident; - at boot it waits up to 2 minutes for yggdrasil to report our address;
- summaries are cached for
cache_seconds(30) and computed one at a time, so repeated requests don't load a small VPS; - re-running
install.shwith new--allowvalues updates only thewebsection of an existing config and regenerates the nft rule.
Find a device's ygg address with yggdrasilctl getSelf.
What's on the dashboard:
- current mode and the share of time spent in each mode;
- node cards: whether packets get through, RTT, loss, and availability in "open" and "whitelist" mode separately;
- timeline: a mode strip + a strip per check (green / partial / none);
- ping6 RTT to both nodes, with whitelist periods shaded;
- "availability by network mode" table: every layer × every mode, the main answer;
- peers now (uptime, latency, traffic, last error) and whitelist episodes with node availability inside each one.
Data and SQL
/var/lib/ygg-mon/ygg-mon.db, tables runs, targets, checks, annotations, kv,
plus the v_checks view that joins them. About 1 MB per day, kept for
retention_days (60).
-- node availability by mode over the last week
SELECT target, mode, COUNT(*) n, ROUND(100.0*SUM(ok)/COUNT(*),1) pct, ROUND(AVG(rtt_ms),1) rtt
FROM v_checks WHERE layer='node' AND ts > strftime('%s','now','-7 days')
GROUP BY target, mode;
-- why peers failed to come up under whitelists
SELECT target, detail, COUNT(*) FROM v_checks
WHERE layer IN ('underlay','peer') AND mode='whitelist' AND ok=0
GROUP BY target, detail ORDER BY 3 DESC;
Multiple probes
Each probe has its own database and name (probe). To see them in one dashboard,
collect the databases on one machine:
scp probe-lte:/var/lib/ygg-mon/ygg-mon.db /tmp/lte.db
ygg-mon import /tmp/lte.db # idempotent, duplicates are skipped
A probe selector then appears on the dashboard.
Config
/etc/ygg-mon/config.json, see config.example.json. The main keys:
nodes[].host: the node's ygg name; the last good address is cached, and you can setaddrby hand as a fallback;nodes[].tcp_ports: ports for the TCP check through the overlay (an RST also counts as ok);canaries.ru/canaries.foreign: how the mode is detected. Prefer sites that answer from any network: some (gosuslugi.ru, for example) filter traffic and may not answer from every network;ip_family:auto(default) tries IPv4 only when the host has no global IPv6 route, otherwise alternates IPv6/IPv4;4or6forces one family for canaries and underlay;ygg_endpoint: ifyggdrasilctlcan't find the socket by itself (unix:///var/run/yggdrasil/yggdrasil.sock);web.listen: an address,"ygg"(our own ygg address) or a list, e.g.["127.0.0.1", "ygg"];web.allow: ygg addresses/prefixes allowed to open the dashboard (empty = anyone who can reach the listen address).
Limitations
- A QUIC peer can't be checked directly with the stdlib; only the ygg session state covers it.
- The checks cover handshakes and ping, not throughput: the TLS "freeze" after
~16 KB that TSPU applies isn't visible this way. It shows up indirectly as a peer
that comes up and drops right away (
last_error, uptime). - Mode detection from canaries is a heuristic. If your carrier's whitelist differs,
adjust
canaries.ru. - macOS: the collector works (
ping6), but there's no installer; runygg_mon.py collectfrom launchd/cron and pass--config.
License
MIT, see LICENSE.