- Python 99.3%
- Shell 0.6%
Since core 2.61 the SMTP relay is no longer chosen via configured_addr: the core tries the newest relay first and falls back to the next one if a relay is unreachable. The MSG_FAILED failover handler and resilient send patch switched configured_addr and resent, which no longer changes the relay and only produced delayed duplicate resends. - remove on_msg_failed failover and _setup_resilient_mode - /setprimary and /resilient reply with a deprecation note - /transports lists relays in the core's sending order - require deltachat-rpc-server>=2.62.0 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> |
||
|---|---|---|
| .github/workflows | ||
| tests | ||
| .dockerignore | ||
| .gitignore | ||
| bot.py | ||
| CHANGELOG.md | ||
| database.py | ||
| docker-compose.yml | ||
| Dockerfile | ||
| icon.jpg | ||
| icon.png | ||
| README.md | ||
| requirements.txt | ||
| set_admin.py | ||
| update.sh | ||
Delta Chat Uptime Bot
Delta Chat Uptime Bot is a self-hosted uptime monitoring bot (similar to Uptime Kuma) integrated directly with Delta Chat. It monitors resources (websites, APIs, TCP ports, or ping targets) and alerts you inside Delta Chat if they go offline.
Additionally, it automatically generates a secure, beautiful web status dashboard for each chat.
Features
- 🛡️ Secure Administration: Claim ownership with
/initadminvia private chat. Cryptographic fingerprint-based authentication protects administrative actions. - 💬 Per-Chat Isolation: Each chat (private or group) maintains its own separate list of monitored resources.
- 🚨 Incident-Based Alerting & In-Place Dynamic Updates:
- Instead of flooding the chat with dozens of separate DOWN/UP messages, outages trigger a unified Incident per chat.
- As multiple monitors fail or recover, the bot edits the same incident message in-place with real-time status and duration metrics.
- Tiered Rate-Limiting: Live duration updates adaptively back off (every 15s in the first minute, 30s during minutes 1–5, 1m up to an hour, and 5m after an hour) while status transitions edit immediately with zero delay.
- When all services recover, the incident message is updated to Resolved with total downtime duration.
- 🔍 Content & Keyword Assertion (Zero-Config + Custom Keywords):
- Zero-Config Error Detection: In the background, automatically scans HTTP responses for silent failure signatures wrapped in
200 OKresponses (e.g. database connection errors, 502/503 wrapped in HTML, Cloudflare error screens). - Custom Keyword Matching: Assert that response body contains specific text (e.g.
/add https://api.site.com Health "status:ok"or/keyword 1 "Welcome").
- Zero-Config Error Detection: In the background, automatically scans HTTP responses for silent failure signatures wrapped in
- ⏸️ Smart Maintenance Windows & Alert Snoozing:
- Pause monitoring and mute outage alerts during planned maintenance without skewing 30-day uptime metrics:
/pause <id|url> [duration](e.g./pause 1 30m,/pause https://example.com 2h). - Supports replying
/pause [duration]directly to incident alerts. - Automatically resumes normal monitoring when the maintenance window expires, or resume early with
/resume.
- Pause monitoring and mute outage alerts during planned maintenance without skewing 30-day uptime metrics:
- ⚡ Universal Latency Measurement & Response Time Tracking:
- Measures probe latency in milliseconds for HTTP/HTTPS, TCP sockets, and ICMP Ping.
- Real-time latency badges displayed on the Web Status Dashboard (
⚡ 124ms) and in/list.
- 📜 Detailed Outage & Incident History:
/events(or/incidents) — View the chat's historical incident log, active outages, and total downtime durations./history [id]— Inspect recent downtime events for a specific monitor with failure reasons, error codes, and recovery timestamps.
- 🌐 Host Outage Protection & Circuit Breaker:
- Automatically verifies host internet connectivity via high-speed canary checks (
1.1.1.1,8.8.8.8,9.9.9.9,1.0.0.1) before declaring any resource DOWN. - If the bot host itself loses internet access, false mass-downtime alerts and false downtime logs are suppressed, keeping 30-day uptime metrics accurate.
- Automatically verifies host internet connectivity via high-speed canary checks (
- 🔒 SSL Certificate Expiration Monitoring: Automatically tracks SSL/TLS certificate expiration for HTTPS targets:
- Periodic checks cached to run at most once per hour.
- Staged proactive alerts sent to chat at 7 days, 3 days, and 24 hours (1 day) before expiration, as well as upon expiration.
- Automatic alert state reset when certificate is renewed.
- Real-time expiration countdown displayed in
/listand on the Web Status Dashboard.
- 🧹 Stale Resource Notices & 30-Day Auto-Cleanup:
- 7-Day Notice: Sends a notice when a resource is continuously unreachable for 7 days, suggesting removal if decommissioned.
- 14-Day Warning: Sends a warning at 14 days of continuous downtime, advising that 30-day unreachable monitors are automatically removed.
- 30-Day Auto-Cleanup: Automatically deletes resources with continuous 0% uptime for 30 days and notifies the chat of the removal.
- 🛰️ Distributed Multi-Node Peering & Cross-Probe Verification:
- Link multiple bot instances across different regions/servers as remote probes via private 1:1 Delta Chat DMs (
/addpeer <email> [node_name]). - Zero Group Spam: All protocol handshakes, background telemetry, and instant cross-checks happen in private DMs between bots.
- Cross-Probe Verification: Outages are verified across remote probes in real-time before alerting, distinguishing global downtime from regional/routing reachability issues.
- Aggregated Web Dashboard: Web status pages show latency and status badges for all active probe locations (
[📍 Frankfurt-DE: 18ms] [🛰️ RU-Moscow: 45ms]).
- Link multiple bot instances across different regions/servers as remote probes via private 1:1 Delta Chat DMs (
- 🤖 Identified User-Agent: Sends a custom
User-Agentheader (e.g.DeltaChat-Uptime-Bot/2.9.0 (https://git.gluek.info/gluek/deltachat_uptime)) during HTTP checks so server administrators can easily identify monitoring requests in server logs. - 🚀 High-Concurrency Scaling Architecture:
- Two-Tier HTTP Probing (
HEAD->GETwith Auto-Memorization): Zero body download (0 bytes) and 0 CPU decoding for standard HTTP checks viaHEAD, with automatic 16 KBGETfallback on 405 or errors, and 128 KB limit for keyword assertions. When a server responds with405 Method Not AllowedtoHEAD(e.g. GoToSocial, Mastodon, specialized APIs), the bot automatically memorizesGETfor that monitor and database record so all subsequent checks useGETdirectly without generating repeated405log entries. - Async DNS Resolution & 5m TTL Cache:
aiodns(AsyncResolver) support with 300s TTL cache prevents blocking DNS lookups and reduces repeated nameserver queries. - Native Async ICMP Ping (
aioping): Native in-process raw socket pinging eliminates subprocess creation overhead, with automatic fallback to/bin/pingif raw socket permissions are restricted. - Deterministic Time Slot Staggering: Uniform phase slotting (
(r_id * 11) % interval) spreads checks across 5-second windows, preventing thundering-herd spikes on startup and interval boundaries. - Lock-Free SQLite WAL Concurrent Reads & Persistent Writer: Dedicated
_write_lockprotects write transactions with a persistent connection (_writer_conn) while enabling concurrent lock-free reads for web dashboards and chat status commands. - Zero-Churn Disk I/O (
synchronous=NORMAL): Uses SQLite WAL mode withPRAGMA synchronous = NORMAL, eliminating fsync overhead during commits and reducing continuous disk writes by >98%. - Non-Blocking Semaphore: Concurrency semaphore slots are held exclusively for the milliseconds of network probes; retry backoffs (30s) and remote peer checks execute asynchronously without starving healthy checks.
- Dedicated Thread Pools: Separate thread executors for database queries (
uptime_db) and Delta Chat JSON-RPC / SMTP calls (uptime_rpc) prevent slow email delivery from stalling database operations. - Single-Query Batch Metrics & In-Memory TTL Cache: Single SQL batch queries and 60-second TTL caching eliminate N+1 queries across web status pages,
/list, and/status. - Optimized SQLite Indexes: Comprehensive indexes on downtime intervals, incident states, and peer telemetry ensure sub-millisecond query execution even with tens of thousands of historical records.
- Two-Tier HTTP Probing (
- 🔄 Failure Resiliency & Retry Logic:
- Checks resources once a minute.
- If a resource check fails, the bot does not alert immediately. It retries 2 more times at 30-second intervals.
- Alerts are only triggered if all 3 checks fail, avoiding false positives.
- Once a DOWN resource recovers, it is marked UP on the first successful check.
- 📊 Uptime Dashboards: Generates a secure, 12-character unguessable base62 URL (e.g.
https://up.example.com/k8D2x9mPqL1a) hosting a modern dark-themed web status dashboard with active status, latency metrics, SSL countdowns, and recent incident logs for each chat. - ✉️ Multiple Mail Relays: Supports multiple mail servers. Relay selection and failover are handled by the Delta Chat core (2.61+), which sends via the newest relay first and falls back to the next one if a relay is unreachable.
Commands
User Commands (Public per-chat)
These commands are available to any member of a chat. They support suffixes (e.g. /add@up, /status@uptime) to route commands correctly if multiple bots exist in the same chat.
A plain /help sent in a group chat is answered in a private 1:1 chat with the sender, so several bots don't flood the group with help texts. Use /help@uptime to show the help in the group itself.
/ping <target> ["keyword"](or/check,/test) — Perform an immediate on-demand reachability and diagnostics test without adding to monitoring (rate-limited to 1 request per 15s per chat for non-admins; supports HTTP/HTTPS, SSL certificate inspection, keyword assertions, TCP port, ICMP Ping, and multi-region cross-checks)./add <target> [name] ["keyword"]— Add a monitor. Target formats: •https://google.com Google(HTTP/HTTPS check) •https://api.site.com Health "status:ok"(HTTP with keyword assertion) •google.com:443 Google TCP(TCP port check) •google.com Google Ping(ICMP Ping check)/remove <id|url>(or/delete,/rm,/del) — Stop monitoring a resource by ID or target URL. You can also reply/removedirectly to any incident or outage alert message to remove the affected monitor without knowing its ID./pause <id|url> [dur](or/mute,/maintenance) — Mute outage alerts during maintenance (e.g./pause 1 30m,/pause https://example.com 2h, or reply/pauseto an alert)./resume <id|url>(or/unpause) — Resume active monitoring after maintenance./keyword <id|url> [keyword|none]— Configure or remove expected keyword assertion./list— List monitored resources, status, latency, and SSL expiration in this chat./status— View monthly uptime statistics and get the link to the chat's secure Web Status Page./events— View recent incidents and active outages for this chat./history [id]— View downtime history for monitors./sync— Synchronize monitored resources with other bots in the same chat (rate-limited to 1/minute for non-admins)./donate— Support bot development ❤️/help— View available commands and system information.
Admin-Only Commands
These commands are only executable by the configured administrator.
/url— View current base external status URL./url <url>— Update the base external status URL (e.g.,/url https://up.gluek.info) to generate correct status links./nodename [name]— View or set the local probe node identifier (e.g./nodename Frankfurt-DE)./invitepeer— Generate a SecureJoin E2E encrypted invite link (https://i.delta.chat/#...) for pairing with other bots./peers(or/probes) — List distributed monitoring peers, node names, and last seen activity./addpeer <email|link> [node_name]— Link another Delta Chat Uptime bot as a remote probe using its email address or SecureJoin invite link (e.g./addpeer https://i.delta.chat/#... RUor/addpeer ruptimebot@chat.gluek.info RU)./rmpeer <email>— Remove a remote peer probe./accounts— List active bot accounts./rmaccount <id>— Delete a bot account./transports— Show configured mail relays, status, and stats./addtransport— Add backup mail relays (either chatmail URIs or address/password). Restricted to private 1:1 chat with the bot for security./rmtransport <addr>— Remove backup mail relay.
Relay selection and failover are handled by the Delta Chat core (2.61+): it sends via the newest relay first and falls back to the next one if a relay is unreachable. /transports lists relays in that order. The former /setprimary and /resilient commands are deprecated and only reply with this explanation.
Deployment
Prerequisites
- Docker and Docker Compose installed.
- A dedicated email address for the bot (e.g.
uptimebot@yourdomain.com). - A domain name pointing to your host for the status pages (e.g.
up.gluek.info).
1. Build and Prepare
Clone the repository, enter the directory, and build the Docker container:
cd deltachat_uptime
docker compose build
2. Configure Email and Admin
Initialize the bot account with your email and password:
docker compose run --rm uptime_bot python bot.py init uptimebot@yourdomain.com "your_email_password"
Configure your admin email address and optionally your cryptographic fingerprint on the server:
docker compose run --rm uptime_bot python set_admin.py --email admin@yourdomain.com
docker compose run --rm uptime_bot python set_admin.py --fingerprint 1234ABCD1234ABCD1234ABCD1234ABCD1234ABCD
3. Run the Bot
Start the bot daemon in background:
docker compose up -d
4. Claim Ownership inside Delta Chat
- Scan the bot's secure join QR code printed in the logs (
docker compose logs uptime_bot), open the root web dashboard (/) to scan the high-resolution vector QR code, or copy the invite link / email address into Delta Chat. - Send
/initadminto the bot in a private message. - The bot will automatically verify your identity and associate your cryptographic fingerprint. You are now the administrator!
5. Set up Base URL
Tell the bot your public status domain so that it generates correct dashboard links:
/url https://up.gluek.info
Configuration & Profile Customization
You can customize the bot's name, avatar, and status text by passing environment variables (e.g. in your .env file or docker-compose.yml):
DISPLAY_NAME— Customize the display name of the bot (default:Delta Chat Uptime Bot).STATUS_TEXT— Customize the status/about text of the bot.AVATAR_PATH— Path to an image file (PNG/JPG) to use as the bot's profile avatar (default: falls back toicon.pngoricon.jpgin the project root).
Reverse Proxy with Caddy
If you use Caddy on your host (like for your ntfy bot), you can expose the status pages by adding the following config to your /etc/caddy/Caddyfile:
up.gluek.info {
reverse_proxy 127.0.0.1:8080
}
Reload Caddy to apply changes:
sudo systemctl reload caddy