Software Engineer
Loading posts...
The full service map of my homelab: six hosts, ~100 containers, every stack and the wiring between them. The deep-dive companion to the overview post.

I put an interactive industrial robot arm on my homepage. Here is the short version of why, and the parts that turned out to be more interesting than I expected.
Feel free to contact me at kanishksachdev@gmail.com
Six hosts, about a hundred containers, one Tailscale mesh, and a Mac that runs deploys but holds no state. This post is the reference diagram set for the whole thing; I rewrite it when the fleet changes shape. Last major revision: August 2026.
| Metric | Count |
|---|---|
| Hosts | 6 (3 cloud VPSes, 3 in my apartment) |
| Services defined | 102 fleet-wide, 68 on the main box |
| Public DNS records | 48, managed as code with DNSControl |
| Origin TLS certificates | 38 (Let's Encrypt via traefik) |
| SSO providers | 17 (9 proxy forward-auth, 8 native OIDC), all in YAML blueprints |
| Git repos | 1 parent + 8 submodules, every host is its own repo |
| Compose files on the main host | 15, merged with include: |
| Ansible playbooks / roles | 14 / 5 |
| Backup cadence | hourly borg, keeps 24 hourly / 7 daily / 4 weekly / 6 monthly |
| Secret scanning | gitleaks pre-commit hook + CI in all 9 repos |
| Time to ship a new SSO-gated app | ~10 minutes, three file edits |
| Cost | ~$75/mo, breakdown below |
Loading diagram…
| Host | Hardware | Role |
|---|---|---|
| hetzner | 16 vCPU / 32 GB VPS, 148 GB attached volume as docker data-root | The main box: ingress, auth, mail, media, photos, finance, secrets, observability. Only host with public ports. |
| oci | 4-core Ampere ARM64, 24 GB, Oracle free tier | Secondary DNS (Technitium) plus the standard agent set |
| transmission | 1 vCPU / 2 GB | Torrent client behind a Mullvad WireGuard tunnel, Sweden exit. Isolated on purpose. |
| rpi | Raspberry Pi 4B, 8 GB, on apartment WiFi | Home Assistant, proxied out through hetzner over Tailscale |
| optiplex | Dell OptiPlex 7040, 8 threads / 31 GB, residential connection | Jobs that need residential egress: some sites challenge every datacenter IP they see |
| kanishk-desktop | A laptop that is asleep more often than not | PhotoPrism for a 24 GB library, monitoring agents. Ansible treats "unreachable" as normal for this one. |
The three home hosts are the newest additions. The Pi is the second board to hold that name after the first one died; the laptop deploys like every other host but is also reachable as a headless desktop through Authentik's RAC outpost, which hands you a GNOME session over RDP in a browser tab.
One parent repo, eight submodules. Six of them are host configs, and each is
also the working tree on its host: /root on hetzner is a checkout of
hetzner-vps-config, and a deploy is a git push from the Mac followed by an
Ansible-driven git pull on the host. The other two are dns/ (the public
zone as code) and docs/ (private runbooks). The Mac triggers deploys but is
never the source of file content. If it's not committed, it doesn't exist.
Loading diagram…
Renovate is self-hosted and covers all nine repos through a GitHub App, with
every image pinned to tag + digest. When it bumps a version, the merge is the
deploy decision and deploy.yml is the deploy.
Every public request crosses five layers before it reaches application code.
Loading diagram…
Some detail that the diagram flattens:
Around 30 web apps hang off *.kanishksachdev.com, all through the one
traefik. Grouped, with how each is gated:
| Group | Hostnames | Access |
|---|---|---|
| Identity | auth | Authentik itself |
| Infra admin | traefik crowdsec files houndarr portainer | forward-auth, admin group only |
| Observability | grafana status | Grafana via OIDC; the status page is public on purpose |
| Media | jellyfin requests sonarr radarr bazarr prowlarr transmission dispatcharr | Jellyfin native OIDC; the arrs behind forward-auth with API-path bypass routers so their own API keys still work |
| Photos | grad photos | one public gallery, one gated instance served off the laptop |
mail email ntfy | mailserver accounts; ntfy has its own ACL | |
| Personal apps | sure actual karakeep closet ha | native OIDC where the app supports it |
| Dev / misc | secrets renovate dns pocketbase centrifugo home minecraft | own auth; minecraft is a raw TCP port |
The forward-auth pattern costs three file edits per app: a compose entry with traefik labels, an Authentik blueprint (provider + application + policy binding), and a line in the outpost's provider list. Push, deploy, done, certificate included.
Two auth flows coexist:
Loading diagram…
Everything in Authentik lives in 25 blueprint YAML files in git: 17 providers, applications, policy bindings, groups, and the recovery and enrollment flows. Secrets in blueprints resolve from the environment at apply time, so the files are committable. No clicking around in an admin UI for permanent state.
Loading diagram…
Bitmagnet deserves its own paragraph. It crawls the BitTorrent DHT and exposes what it finds to Prowlarr as a Torznab indexer: no tracker sites, nothing to Cloudflare-challenge. The catch: it ingests roughly 413k torrents a day and only ~7.5% are the movie/TV content the service exists for. Left alone it grew the database 3.8 GB per day. Taming it took two layers: classifier flags that delete unwanted categories at ingest, and a nightly timer that ages out never-classified rows, with dead-man's-switch metrics so Prometheus alerts if the pruning silently stops. Disk growth is now flat.
The public zone lives in a dnsconfig.js and applies to Cloudflare through
DNSControl. ./dns.sh preview diffs config against live before every push,
and CI runs a daily drift check that pings ntfy when reality stops matching
git. That catches both my dashboard shortcuts and anything else that edits
the zone behind my back. Applying is deliberately manual: a DNS diff is worth
reading every time, because per-record proxy flags are doing real security
work and a dropped one fails silently.
Self-hosted Infisical is the source of truth. Each host has a prefix
(HETZNER_*, OCI_*, ...) and a read-only machine identity; the deploy
renders matching secrets into the host's .env before compose runs, so adding
a secret is: put it in Infisical, reference ${VAR} in compose, deploy. The
one bootstrap credential that can't live in Infisical, the identity that
unlocks it, is ansible-vault encrypted, with the passphrase kept outside the
worktree where git cannot see it.
The enforcement half: a gitleaks pre-commit hook runs on every commit in every repo via a global hooks path, and a gitleaks workflow scans full history in CI on every push. Allowlists are committed per-repo, scoped by path and regex, so public-by-design values like OAuth client IDs pass while an actual client secret in a staged diff blocks the commit. I also audited all nine repos' full git history against every credential the fleet has ever used and rewrote where needed. That was an afternoon I'd rather not repeat, hence the hook.
Loading diagram…
All 66 containers on the main host stream logs to Loki, which clusters incoming lines into pattern templates: a wall of near-identical log lines collapses into a handful of shapes with placeholders where the variables were. Uptime checks live on a public Kener status page, and the probes against SSO-gated apps authenticate with a service-account app password, because an unauthenticated check happily follows the redirect to the login page, gets a 200, and reports a dead service as up.
Borg runs hourly on the main host over SSH to the storage box, with 24 hourly / 7 daily / 4 weekly / 6 monthly retention and a weekly integrity check. The check is scheduled at 04:30, not 04:00. A borg repo takes an exclusive lock, and a check at the top of the hour raced the hourly backup so one of the two died every Sunday. Which one lost alternated week to week, making it look like random flakiness instead of a standing bug. Both jobs now also wait up to 15 minutes for the lock instead of borg's default of about a second. Failures with exit ≥ 2 page me through ntfy; exit 1 is borg saying "a file changed while I read it," which is background noise on a live system.
The current shape of the fleet is mostly scar tissue. The instructive ones:
The deploy that broke traefik and then couldn't fix it. A commit moved a
bouncer API key from a committed file into a secret-rendered one. The deploy
sequence is pull → render → compose, and traefik's file provider reloads on
the pull, so for a moment it referenced a file that didn't exist yet, the
middleware failed closed on the entrypoint, and every site returned 404.
Re-running the deploy couldn't fix it: the render step fetched secrets through
the very proxy that was down. Recovery came from git show HEAD~1. Two rules
fell out: config that references a rendered file ships in a separate deploy
after the file exists, and the secrets API is now reachable over the tailnet
directly, so the control plane no longer depends on the data plane it repairs.
The bouncer that took the fleet down by restarting. In its default live mode the CrowdSec traefik plugin checks every request against the CrowdSec API and fails closed. CrowdSec restarts on any deploy touching its config, so one bad config crash-looped it and every hostname served 403 for five minutes. Stream mode inverts the failure: the plugin caches the ban list and a dead CrowdSec degrades to slightly stale bans instead of a total outage. I verified by stopping CrowdSec outright and watching everything keep serving.
The alert channel that noise silently killed. Traefik's ACME store keeps certificates for services that no longer exist, they get staler forever, and each one fires an expiry alert indefinitely. Eight dead certs accumulated; alertmanager groups by alertname, so all of them rendered into one ntfy message, which then exceeded ntfy's size limit and was rejected outright, not truncated. Past a threshold, alert noise stops being a nuisance and becomes an outage of the alerting itself. The store got surgically cleaned and the alert list went from 89% noise to quiet.
DKIM that permfailed for months without a bounce. The published DKIM
record had been silently mangled in a dashboard paste. Every signature
failed, but SPF alignment carried DMARC, so mail flowed and nothing
alerted. Found only by actually running the verifier against the live zone.
The mail stack's DNS posture (SPF -all, DMARC p=reject, 2048-bit DKIM)
gets checked from inside the container now, not assumed from the config.
| If this dies | Gone | Still standing |
|---|---|---|
| hetzner | Everything public: auth, mail, media, all routing | The other five hosts keep running their local stacks |
| oci | Secondary DNS, its monitoring agents | Everything user-facing |
| transmission | New downloads queue up | Streaming existing media |
| rpi | Home Assistant | Everything else |
| optiplex | Residential-egress jobs | Everything else |
| kanishk-desktop | Nothing anyone would notice; it's usually asleep | Everything |
| Storage box | Media library, photo originals, and the borg target. One failure domain, top of the fix list | Apps that don't touch bulk storage |
| Cloudflare | The proxied edge | Origin still answers |
| GitHub | New deploys and Renovate PRs | Everything already running |
| Infisical | The next deploy's secret render | Running services keep their env |
| Authentik | Every gated app | Apps with their own auth |
| The Mac | My ability to deploy until I reach another machine | The entire fleet; nothing runs on the Mac |
| Item | ~$/mo |
|---|---|
| Main VPS (16 vCPU / 32 GB) | 38 |
| 148 GB attached volume | 8 |
| Storage box (5 TB) | 13 |
| Transmission VPS | 5 |
| Oracle free tier | 0 |
| Tailscale, Cloudflare, Let's Encrypt | 0 |
| Total | ~$64 |
Everything else (SSO, mail, secrets manager, dependency automation, status page, monitoring, photo hosting, bookmarks, finance) is self-hosted on the boxes above. The managed-service equivalent runs several times that, but the honest accounting is that I run this because I want to, and the skills are the dividend.
root. Renaming it
would orphan every named volume. Same reason two finance containers keep
names that read backwards: renaming the services would orphan the volumes
holding the actual financial data..gitignore in a host repo is a whitelist: ignore everything, then
un-ignore what's deliberately tracked. New files default to "not committed,"
which is the right default when the repo is also a home directory. The
sharp edge: the whitelist swallows new dotfiles and directories silently,
so .github/ and friends need explicit un-ignore lines.| Layer | Tool |
|---|---|
| Reverse proxy | Traefik v3, single ingress |
| Edge protection | Cloudflare proxy + CrowdSec (HTTP + nftables bouncers) |
| Auth | Authentik, blueprint-managed, forward-auth + OIDC |
| Secrets | Infisical, self-hosted, per-host machine identities |
| Secret scanning | gitleaks: pre-commit hook + CI, all repos |
| Orchestration | Docker Compose + Ansible from a Mac |
| Dependency updates | Renovate CE, self-hosted, digest pinning |
| DNS | Cloudflare zone via DNSControl; Technitium internally |
| Metrics | Prometheus + Alertmanager → ntfy |
| Logs | Loki + promtail, pattern detection on |
| Dashboards | Grafana; Homepage as the front door; Kener for public status |
| Backups | Borg over SSH, hourly |
| docker-mailserver + Roundcube, DKIM/SPF/DMARC all strict | |
| Media | Jellyfin + the arr suite + bitmagnet + dispatcharr |
| Storage | Hetzner Storage Box, CIFS, 5 TB |