[{"data":1,"prerenderedAt":452},["ShallowReactive",2],{"category:build-log":3},{"category":4,"index":10,"total":14,"others":15,"posts":61},{"id":5,"description":6,"extension":7,"label":8,"meta":9,"order":10,"post_count":10,"slug":11,"stem":12,"__hash__":13},"categories\u002Fcategories\u002Fbuild-log.json","Field notes from building specific systems, in progress.","json","Build Log",{},2,"build-log","categories\u002Fbuild-log","_YPWw94FFI3NAi5oP4e4Z_N0AcFBLoywskvd-MT0les",6,[16,26,35,43,52],{"id":17,"description":18,"extension":7,"label":19,"meta":20,"order":21,"post_count":22,"slug":23,"stem":24,"__hash__":25},"categories\u002Fcategories\u002Fcraft.json","Close-read posts on the daily work of shipping software.","Craft",{},0,1,"craft","categories\u002Fcraft","bAfaerFLn5gKJyzXQLTEy95inz6saFE03vKNGDNWwNw",{"id":27,"description":28,"extension":7,"label":29,"meta":30,"order":22,"post_count":31,"slug":32,"stem":33,"__hash__":34},"categories\u002Fcategories\u002Fskeptics-toolbox.json","The cases against default narratives; what doesn't hold up under scrutiny.","Skeptic's Toolbox",{},3,"skeptics-toolbox","categories\u002Fskeptics-toolbox","hPOh22rNWxW_dewYrYcLgnrEDIv_sY17XJAlLjRxTe4",{"id":36,"description":37,"extension":7,"label":38,"meta":39,"order":31,"post_count":22,"slug":40,"stem":41,"__hash__":42},"categories\u002Fcategories\u002Fml-journal.json","Learning ML in public. Failures, diagnoses, partial wins.","ML Journal",{},"ml-journal","categories\u002Fml-journal","0PP-FTOq2_F3lMsfbjdhi3q9TqLm3ajKZCqt9kXAXCM",{"id":44,"description":45,"extension":7,"label":46,"meta":47,"order":48,"post_count":22,"slug":49,"stem":50,"__hash__":51},"categories\u002Fcategories\u002Fknowledge-architecture.json","How knowledge is structured, captured, retrieved, transmitted.","Knowledge Architecture",{},4,"knowledge-architecture","categories\u002Fknowledge-architecture","acvknIIMKZMnFb9-VBtiSIDksRlzGbGMQUpyFX-SMns",{"id":53,"description":54,"extension":7,"label":55,"meta":56,"order":57,"post_count":22,"slug":58,"stem":59,"__hash__":60},"categories\u002Fcategories\u002Fpeople.json","Stories about humans in the work — collaborators, clients, collaborators in absentia.","People",{},5,"people","categories\u002Fpeople","9rpzEOrRX1WY5EmvqmKGOyGuAwvs2IXaPT9IfCoJfn4",[62,328],{"id":63,"title":64,"body":65,"category_label":8,"category_slug":11,"date_created":315,"dek":316,"description":307,"directus_id":317,"extension":318,"is_promoted":319,"meta":320,"navigation":321,"path":322,"published_at":315,"seo":323,"slug":324,"status":325,"stem":326,"__hash__":327},"posts\u002Fposts\u002Fquarterly-k8s-migration-is-a-tuesday-night.md","The quarterly K8s migration is now a Tuesday night",{"type":66,"value":67,"toc":306},"minimark",[68,73,82,85,88,92,95,117,120,124,131,134,138,141,144,204,208,211,272,279,283,290,300,303],[69,70,72],"h2",{"id":71},"the-thesis","The thesis",[74,75,76,77,81],"p",{},"Production K8s migrations (the real ones, with stateful data and live customer traffic) are ",[78,79,80],"strong",{},"evening-tier work in 2026",", not quarter-tier projects. Most engineering orgs haven't updated their mental model. They're still scheduling \"infrastructure modernization initiatives\" with kickoff decks, weekly syncs, and a target completion of \"end of next quarter.\" Then they wonder why nothing actually moves.",[74,83,84],{},"Tonight we moved 38 pods, ~150 GB of state, an entire ES cluster (including its master quorum), and an ingress's worth of traffic across two AZs onto storage that didn't exist on those nodes 36 hours ago. With 13 minutes of outage. From a French ISP-grade home setup, with a CTO who said \"go faster, I'm taking the risk\" and a Claude Code window.",[74,86,87],{},"This is what tooling has done for us. People should know.",[69,89,91],{"id":90},"what-we-actually-did","What we actually did",[74,93,94],{},"Evening of 2026-05-05 → morning of 2026-05-06:",[96,97,98,105,111],"ul",{},[99,100,101,104],"li",{},[78,102,103],{},"Stage 2",", burn-in environment cutover to the new nodes. 9 pods. Validated the storage path and the compute uplift before touching prod.",[99,106,107,110],{},[78,108,109],{},"Stage 3",", production application (the API, the queues, redis, the FTP intake). 23 pods. 4 phases (stateless → redis → import-ftp → re-enable scheduled imports). 3 minutes of customer outage during the redis cutover, the only window where we couldn't roll.",[99,112,113,116],{},[78,114,115],{},"Stage 4",", Elasticsearch. 6 nodes restructured (legacy 3-data + 2-master + 1-transform on Longhorn → 2-data + 2-master + 1-voting_only + 1-transform on LINSTOR with daily S3 snapshots). 10 minutes of customer outage from a stale env-var oversight.",[74,118,119],{},"Replicas, replicated storage, S3 backups, anti-affinity, voting-only tiebreaker on a tiny VM in another AZ. The whole shape of \"how a real prod cluster looks.\" Done in an evening. Walked back into the room the next morning, dashboards green, channel quiet.",[69,121,123],{"id":122},"the-compute-claim","The compute claim",[74,125,126,127,130],{},"Mid-migration, on a real workload (a 63 MB XML import that fans out into 24,000 child jobs across 6 Horizon workers), the new nodes processed it ",[78,128,129],{},"5.7× faster per job"," than the equivalent on legacy. 58 jobs\u002Fsec vs 10 jobs\u002Fsec, same code, same Redis, same Postgres. Just newer silicon and proper local-NVMe-backed DRBD instead of whatever the legacy storage was doing.",[74,132,133],{},"That number alone justifies the migration. Doing it on a Tuesday night justifies stopping the meeting.",[69,135,137],{"id":136},"what-made-it-possible","What made it possible",[74,139,140],{},"It wasn't the execution. The execution was just typing the script. That's the whole point.",[74,142,143],{},"It was:",[145,146,147,153,176,198],"ol",{},[99,148,149,152],{},[78,150,151],{},"Two months of paranoid planning."," Every stage, patch file, and cutover sequence on paper. The outage windows pre-budgeted. The rollback steps documented. By the time we ran a command tonight, we'd already played the move in our heads four or five times. Planning is the bottleneck, not execution. Always was. Tooling just made execution cheap enough that the planning effort actually pays off.",[99,154,155,158,159,163,164,167,168,171,172,175],{},[78,156,157],{},"The runbook structure."," Every destructive step lived in a script under ",[160,161,162],"code",{},"infra\u002Ftmp\u002F"," (one-shot) or ",[160,165,166],{},"infra\u002Fclusters\u002Fprod\u002F"," (canonical \u002F reproducible). Idempotent. Re-runnable. When the first containerd-move script aborted halfway because ",[160,169,170],{},"k8s-agent.service"," didn't exist (it was ",[160,173,174],{},"kubelet.service","), I just edited the failed step and re-ran. Recovery time: 30 seconds. Not \"schedule a follow-up call.\"",[99,177,178,181,182,185,186,189,190,193,194,197],{},[78,179,180],{},"The cutover tricks."," The cleanest one tonight was for the redis migration. Standard playbook says: BGSAVE, dump.rdb out, delete the old PVC, recreate on new storage, copy dump.rdb in, start redis. The race is the gap between ",[160,183,184],{},"kubectl apply"," (which scales redis up) and your manual ",[160,187,188],{},"scale --replicas=0"," (which stops it before it overwrites your dump.rdb with empty data). My CTO suggested overriding the redis container's command to ",[160,191,192],{},"sleep 2d"," for the duration of the data restore. Pod runs, mounts PVC, doesn't run redis, can't write to disk. We ",[160,195,196],{},"kubectl cp"," the data in, then re-apply with the real command. Race eliminated. The kind of move you can only see if you understand both the system AND the orchestrator's idempotency model.",[99,199,200,203],{},[78,201,202],{},"The CTO who took the risk."," Originally this was supposed to be a 48-72-hour burn-in followed by a careful prod cutover spread over a week. He looked at the data after 4 hours and said \"go faster, I'm taking responsibility.\" That is the rarest thing in engineering management: someone who reads the signals and writes \"go\" in the chat. Without that, this is still a quarterly project. With that, it's tonight.",[69,205,207],{"id":206},"what-surprised-us","What surprised us",[74,209,210],{},"The surprises tonight, none of which broke us:",[96,212,213,235,249],{},[99,214,215,218,219,222,223,226,227,230,231,234],{},[78,216,217],{},"DEV1-M comes with a 9 GB root disk, not the 50 GB I'd assumed during planning."," Containerd images alone trip kubelet's ",[160,220,221],{},"DiskPressure"," threshold under default eviction policy. The voting-only ES master couldn't schedule on the tiebreaker node. Fix: attach a 50 GB SBS volume, restructure as ",[160,224,225],{},"\u002Fmnt\u002Fsbs"," with bind mounts to ",[160,228,229],{},"\u002Fvar\u002Flib\u002Fcontainerd"," and ",[160,232,233],{},"\u002Fvar\u002Flog",". Discovered, fixed, documented in a runbook, committed. 90 minutes door-to-door. Three years ago this would have been a Slack thread that died after two days because nobody could agree on the right approach.",[99,236,237,244,245,248],{},[78,238,239,240,243],{},"Bitnami pulled ",[160,241,242],{},"bitnami\u002Fredis:8.0"," from Docker Hub silently sometime in August 2025",", as part of their licensing changes. Our prod redis was running on a cached image. The moment it ever restarted, it was dead. We caught this in Stage 2 burn-in (where the image had to pull fresh on a new node) and switched to ",[160,246,247],{},"bitnamilegacy\u002Fredis:8.0",". It's dumb luck that this didn't blow up production three months ago.",[99,250,251,254,255,258,259,262,263,266,267,271],{},[78,252,253],{},"ECK auto-deletes per-nodeSet headless services when you remove a nodeSet."," Our prod API was hardcoded to ",[160,256,257],{},"search-es-es-data.elasticsearch.svc",", the headless service of the legacy data nodeSet. When Stage 4 removed that nodeSet, ECK deleted the service. The public site's ",[160,260,261],{},"\u002Facheter\u002F*"," routes started returning 500. 10 minutes of customer outage. Switched to the cluster-wide LB service ",[160,264,265],{},"search-es-es-http.elasticsearch.svc",", rolled the API pods, recovered. ",[268,269,270],"em",{},"Never reference per-nodeSet services from app config."," Now permanent in my head, and in the recap doc.",[74,273,274,275,278],{},"These are the kind of things you only learn by doing it. The kind of things people who've never done a real migration don't know they don't know. The kind of things that ",[268,276,277],{},"would"," eat a quarter, if you let them.",[69,280,282],{"id":281},"the-ceiling-moved","The ceiling moved",[74,284,285,286,289],{},"Most engineering orgs are still in the \"infrastructure modernization is a quarterly initiative\" mental model. They have a Confluence page. They have a Q3 OKR. They have a Slack channel called ",[160,287,288],{},"#infra-modernization-2026"," that gets one message a week.",[74,291,292,293,295,296,299],{},"The actual ceiling moved. Container Storage Interface drivers are stable. Operators (ECK, Piraeus, etc.) handle the orchestration that used to need humans (shard drains, voting-config exclusions, PVC binding races). S3 snapshots are free real estate. Kustomize gives you patches that compose. ",[160,294,196],{}," works through a tar pipe and is fast enough for any data volume that fits in your ",[160,297,298],{},"dump.rdb",".",[74,301,302],{},"If you've been running production K8s for three years, you have all the components. You just need the planning discipline and the runbook structure, plus a CTO who reads the signals. Then it's an evening.",[74,304,305],{},"Mine took 13 minutes of outage. Yours might take 30. It's still a Tuesday.",{"title":307,"searchDepth":10,"depth":10,"links":308},"",[309,310,311,312,313,314],{"id":71,"depth":10,"text":72},{"id":90,"depth":10,"text":91},{"id":122,"depth":10,"text":123},{"id":136,"depth":10,"text":137},{"id":206,"depth":10,"text":207},{"id":281,"depth":10,"text":282},"2026-05-06T01:07:20.191Z","38 pods, ~150 GB of state, and an Elasticsearch cluster moved across two AZs in one evening, with 13 minutes of outage. The execution was never the bottleneck. Two months of planning was.","31992263-3014-46c1-b5eb-81697814192e","md",false,{},true,"\u002Fposts\u002Fquarterly-k8s-migration-is-a-tuesday-night",{"title":64,"description":307},"quarterly-k8s-migration-is-a-tuesday-night","published","posts\u002Fquarterly-k8s-migration-is-a-tuesday-night","IYtXuzngSb7PnXU1LdWXE95wdB5sPyjLQyvEedHVaBQ",{"id":329,"title":330,"body":331,"category_label":8,"category_slug":11,"date_created":443,"dek":444,"description":307,"directus_id":445,"extension":318,"is_promoted":319,"meta":446,"navigation":321,"path":447,"published_at":443,"seo":448,"slug":449,"status":325,"stem":450,"__hash__":451},"posts\u002Fposts\u002Fbuilding-my-own-content-pipeline.md","I Built a Content Pipeline Because I Was Tired of Copy-Pasting",{"type":66,"value":332,"toc":437},[333,337,340,343,346,350,356,371,378,389,392,396,403,406,412,415,420,424,431,434],[69,334,336],{"id":335},"the-problem","The problem",[74,338,339],{},"I write things. Blog posts. LinkedIn takes. X threads, Dev.to write-ups, Bluesky one-liners. Every publish was the same ritual. Write in one place, copy-paste to three others, fiddle the tone, lose track of what went where.",[74,341,342],{},"Not hard work. Death-by-a-thousand-cuts work. The kind that stays under the radar until you realize you've been spending an hour after every post on copy-paste hygiene.",[74,344,345],{},"So I built the pipeline I wished I had.",[69,347,349],{"id":348},"what-i-actually-built","What I actually built",[74,351,352,353,299],{},"Local-first. Docker Compose. Directus as the headless CMS, Postgres underneath, n8n for workflow orchestration. Three containers, one ",[160,354,355],{},"docker-compose up",[74,357,358,359,362,363,366,367,370],{},"The capture surface lives in Claude Code. A skill called ",[160,360,361],{},"capture-thought"," watches for moments mid-session (debugging, refactoring, whatever) when I drop a line like \"save this as a writing idea, btw.\" The skill writes a raw markdown file to ",[160,364,365],{},"~\u002FNextcloud\u002FSync\u002Fwriting\u002Fraw\u002F"," and syncs it to Directus with status ",[160,368,369],{},"idea",". The thought is preserved before it evaporates. Cost to me: half a sentence.",[74,372,373,374,377],{},"Then ",[160,375,376],{},"elaborate-post",", a second skill that takes the raw idea and develops it into a draft. Tags. Excerpt. Cards. Body cleaned up.",[74,379,380,381,384,385,388],{},"The interesting part is what happens after. n8n watches for posts ready to syndicate and routes them to platforms based on tags. ",[160,382,383],{},"tech"," goes to the blog, Dev.to, and X. ",[160,386,387],{},"freelance"," hits LinkedIn. An LLM generates platform-specific variants, same core idea, different voice for each audience. It pushes to LinkedIn (3000 char), X (280), Dev.to, the blog, Malt (2000) and Bluesky (300).",[74,390,391],{},"The routing matrix is the design decision I'm proudest of. Tags route to platforms via a relevance score (high \u002F medium \u002F low \u002F excluded). I don't pick platforms per-post. I tag, and the matrix decides. Per-post selection doesn't scale. Once you have 50 posts, every \"where should this go?\" decision becomes friction. Tag-based routing scales linearly with new platforms, not with new posts.",[69,393,395],{"id":394},"why-not-notion-or-buffer","Why not Notion or Buffer",[74,397,398,399,402],{},"Because I wanted the content to live where I control it. Local Postgres, local files as backup, everything versioned. The whole stack portable via ",[160,400,401],{},"pg_dump"," and a Directus schema snapshot.",[74,404,405],{},"There's an old principle I keep coming back to:",[407,408,409],"blockquote",{},[74,410,411],{},"Every post on someone else's platform is renting space. Every post on your own domain is equity.",[74,413,414],{},"Same logic for the pipeline that produces those posts. If Buffer changes their pricing, breaks their LinkedIn integration, or just disappears, my pipeline doesn't notice.",[74,416,417,418,299],{},"The constraint that drove this is 2AM. I get ideas while debugging production, and I need to grab them in three words before they evaporate. Buffer doesn't speak ",[160,419,361],{},[69,421,423],{"id":422},"the-meta-thing","The meta thing",[74,425,426,427,430],{},"I'm capturing this thought ",[268,428,429],{},"using the very pipeline I'm writing about",", which is either poetic or deeply nerdy. Probably both.",[74,432,433],{},"That recursion isn't an accident. The pipeline is opinionated about one thing: capture must be cheap enough that thinking and writing happen in the same gesture. Once you've built that, every thought you have about the pipeline becomes pipeline content. The system feeds itself.",[74,435,436],{},"That's the part Buffer can't sell you. You have to build it.",{"title":307,"searchDepth":10,"depth":10,"links":438},[439,440,441,442],{"id":335,"depth":10,"text":336},{"id":348,"depth":10,"text":349},{"id":394,"depth":10,"text":395},{"id":422,"depth":10,"text":423},"2026-03-02T16:05:05.672Z","I was burning an hour after every post on copy-paste hygiene. So I built the pipeline I wished I had. Directus, n8n, Claude Code, all local-first.","04649e21-8317-4159-9c8d-b612bcd018fb",{},"\u002Fposts\u002Fbuilding-my-own-content-pipeline",{"title":330,"description":307},"building-my-own-content-pipeline","posts\u002Fbuilding-my-own-content-pipeline","GOZxOoo8pBaQskpKcryttGyGwVnm5KEbPMVD8lA8kpw",1786607568637]