Monthly Technical Newsletter

PLATFORM
OPS

Tech with Vishal Abhinav

Engineering knowledge for DevOps engineers, SREs, and Platform builders — with deep dives into Kubernetes, cloud infrastructure, observability, and modern tooling.

0
Issues Published
250+
Subscribers
<1
Years Running
📖 New — Issue #067
platform-ops ~ issue #057
/ glossary search "systemd" 3 matches across 190 terms category services-systemd-boot / glossary open --term "Init" standalone page found — full deep-dive 15 categories · 190 terms indexed Issue #057 — Linux & Unix Glossary 📖
190 Terms 15 Categories Searchable chmod systemd LVM grep DNS
CI/CD & GitOps K8s Error Runbook CrashLoopBackOff Kubernetes Architecture Node NotReady PersistentVolumeClaim StorageClass CSI Drivers StatefulSets OOMKilled Worker Nodes etcd Quorum Control Plane SCC Violations ArgoCD / Flux Prometheus + Grafana Loki / ELK Jaeger Tracing Canary Rollouts Service Mesh eBPF Platform Engineering Observability SRE Practices CI/CD & GitOps K8s Error Runbook CrashLoopBackOff Kubernetes Architecture Node NotReady PersistentVolumeClaim StorageClass CSI Drivers StatefulSets OOMKilled Worker Nodes etcd Quorum Control Plane SCC Violations ArgoCD / Flux Prometheus + Grafana Loki / ELK Jaeger Tracing Canary Rollouts Service Mesh eBPF Platform Engineering Observability SRE Practices

WHAT WE COVER

The 2026 master map — 685 topics, 45 categories, 13 pillars. Search it, filter it, or open a pillar to see exactly what's shipped and what's queued.

685
Topics
45
Categories
154
Live
148
In pipeline
383
Planned
Live — published & linked
In pipeline — next up
Planned — on the backlog
Click any pipeline or planned topic to be told when it lands
platform-ops ~ knowledge-map
tree platform-ops/ --master-map-2026 platform-ops/ ├── foundation/[ 7 live / 10 ] ├── roadmaps/[ 0 live / 18 ] ├── infrastructure/[ 1 live / 38 ] ├── networking/[ 1 live / 31 ] ├── cloud/[ 0 live / 17 ] ├── delivery/[ 12 live / 76 ] ├── kubernetes/[ 58 live / 58 ] ├── reliability/[ 27 live / 116 ] ├── security/[ 0 live / 46 ] ├── data-and-apps/[ 1 live / 60 ] ├── system-design/[ 0 live / 46 ] ├── commands/[ 44 live / 100 ] └── modern-ops/[ 3 live / 69 ] ls -1t published/ ├── service-mesh-operations.html · issue #067 ├── service-mesh-fundamentals.html · issue #066 ├── kubernetes-cluster-operations.html · issue #065 ├── kubernetes-autoscaling.html · issue #064 ├── kubernetes-scheduling.html · issue #063 ├── kubernetes-config-and-access.html · issue #062 ├── kubernetes-workloads.html · issue #061 ├── openshift-operations.html · issue #060 ├── openshift-networking-storage.html · issue #059 ├── openshift-architecture.html · issue #058 ├── linux-unix-glossary.html · issue #057 ├── linux-troubleshooting.html · issue #056 ├── linux-advanced.html · issue #055 ├── linux-fundamentals.html · issue #054 ├── incident-management.html · issue #053 ├── cicd-gitops.html · issue #052 ├── k8-observability.html · issue #051 ├── k8-storage.html · issue #050 ├── k8-error-runbook.html · issue #049 ├── k8-architecture.html · issue #048 └── k8-networking.html · issue #047 13 pillars, 45 categories, 685 topics — 21 deep-dives online 154 topics live · 148 in pipeline · 383 planned
Every category has its own page — architecture diagram, the issues that cover it, and the full topic list. Browse all 45 →
Open the Roadmaps hub — architecture, issues, full topic list →
Planned
Open the IT Infrastructure hub — architecture, issues, full topic list →
In pipeline
Planned
Open the Storage hub — architecture, issues, full topic list →
Live now
LVM
In pipeline
Planned
Open the Virtualization hub — architecture, issues, full topic list →
In pipeline
Planned
Open the Networking hub — architecture, issues, full topic list →
Live now
DNS
In pipeline
Planned
Open the Cloud hub — architecture, issues, full topic list →
In pipeline
Planned
Open the DevOps hub — architecture, issues, full topic list →
Live now
CI/CD Continuous Integration Continuous Delivery Continuous Deployment Jenkins GitLab CI/CD Deployment Strategies Blue/Green Deployment Canary Deployment
In pipeline
Planned
Open the Containers hub — architecture, issues, full topic list →
In pipeline
Planned
Open the Infrastructure as Code hub — architecture, issues, full topic list →
In pipeline
Planned
Open the Configuration Management hub — architecture, issues, full topic list →
In pipeline
Planned
Open the GitOps hub — architecture, issues, full topic list →
Live now
GitOps Argo CD Flux
In pipeline
Planned
Open the Platform Engineering hub — architecture, issues, full topic list →
In pipeline
Planned
Open the SRE hub — architecture, issues, full topic list →
Live now
MTTR Incident Management RCA Postmortem
In pipeline
Planned
Open the Observability hub — architecture, issues, full topic list →
Live now
Observability Monitoring Metrics Logs Traces Prometheus Grafana Alertmanager Loki Jaeger Distributed Tracing
In pipeline
Planned
Open the Logging hub — architecture, issues, full topic list →
Live now
Linux Logging Journald
In pipeline
Planned
Open the Performance Engineering hub — architecture, issues, full topic list →
Live now
CPU Performance Memory Performance Disk I/O Performance Troubleshooting
In pipeline
Planned
Open the Troubleshooting hub — architecture, issues, full topic list →
Live now
Linux Troubleshooting Kubernetes Troubleshooting Performance Troubleshooting Production Incident Troubleshooting Root Cause Analysis
In pipeline
Planned
Open the Backup & DR hub — architecture, issues, full topic list →
In pipeline
Planned
Open the ITSM & Operations hub — architecture, issues, full topic list →
Live now
Incident Management
In pipeline
Planned
Open the Security hub — architecture, issues, full topic list →
In pipeline
Planned
Open the Supply Chain Security hub — architecture, issues, full topic list →
In pipeline
Planned
Open the Database hub — architecture, issues, full topic list →
In pipeline
Planned
Open the Middleware hub — architecture, issues, full topic list →
In pipeline
Planned
Open the API & Microservices hub — architecture, issues, full topic list →
In pipeline
Planned
Open the Distributed Systems hub — architecture, issues, full topic list →
Live now
Quorum
In pipeline
Planned
Open the System Design hub — architecture, issues, full topic list →
Planned
Open the systemd Commands hub — architecture, issues, full topic list →
In pipeline
Planned
Open the Ansible Commands hub — architecture, issues, full topic list →
In pipeline
Planned
Open the Terraform Commands hub — architecture, issues, full topic list →
In pipeline
Planned
Open the Performance Commands hub — architecture, issues, full topic list →
In pipeline
Planned
Open the AI Infrastructure hub — architecture, issues, full topic list →
In pipeline
Planned
Open the AIOps hub — architecture, issues, full topic list →
In pipeline
Planned
Open the FinOps hub — architecture, issues, full topic list →
In pipeline
Planned
Open the Automation hub — architecture, issues, full topic list →
Live now
Shell Automation Bash Automation Python Automation
In pipeline
Planned
Open the Architecture hub — architecture, issues, full topic list →
In pipeline
Planned

LATEST ISSUES

Every issue is a deep-dive with architecture diagrams, code samples, and battle-tested insights.

New Issue
September 2026 · #067
Service Mesh Operations
mTLS identity that turns policy into a statement about services rather than subnets, the AuthorizationPolicy default that flips one workload to deny while its neighbours stay open, retries that multiply through a call graph, and the trace headers your application still has to forward itself.
mTLS Istio AuthorizationPolicy Retries Multi-Cluster Tracing
Read Now →
67
Previous Issue
September 2026 · #066
Service Mesh Fundamentals
What a sidecar mesh actually buys and what it costs, the four Envoy objects every mesh CRD renders into, what service discovery adds on top of kube-dns, and the Service port name that silently disables every L7 feature.
Envoy Istio Linkerd xDS Discovery
Read Now →
66
September 2026 · #065
Kubernetes Cluster Operations
Pod Security rolled out with warn before enforce, the removed API that takes workloads with it on upgrade, and the difference between an etcd snapshot and a real backup.
Upgrades etcd Velero
65
September 2026 · #064
Kubernetes Autoscaling
The HPA algorithm in one line, the missing resource request that silently disables it, why VPA and HPA fight on CPU, and the pod that pins a node against scale-down.
HPA VPA Cluster Autoscaler
64
September 2026 · #063
Kubernetes Scheduling
Taints as the node's veto, affinity as the pod's request, the topologyKey that gives anti-affinity its meaning, and why Pending is a filtering result you can read verbatim.
Taints Affinity Topology Spread
63
September 2026 · #062
Kubernetes Config & Access
ConfigMaps that update in place except when they do not, Secrets that are encoded rather than encrypted, and RBAC's one rule — purely additive, no deny.
RBAC Secrets Namespaces
62
September 2026 · #061
Kubernetes Workloads
ReplicaSets and why a stuck rollout is legible, DaemonSets that silently never roll, CronJobs that pile up, and the liveness probe that turns a blip into a restart storm.
ReplicaSets Probes CronJobs
61
September 2026 · #060
OpenShift Operations
Monitoring that is half switched off by default, LogQL that returns before it times out, the four security gates, and the upgrade that stops on one PodDisruptionBudget.
Monitoring Loki Upgrades
60
September 2026 · #059
OpenShift Networking & Storage
Routes and the four TLS modes, the OVN MTU fault everybody misdiagnoses, and the CSI chain where each stage fails in a different log.
Routes OVN CSI
59
September 2026 · #058
OpenShift Architecture & Fundamentals
The operator ownership chain under CVO and OLM, projects and the project template, MachineConfig, SCC admission, image streams and etcd — each with the failure it produces.
Operators SCC etcd
58
September 2026 · #057
Linux & Unix Glossary
190 terms across 15 categories — fundamentals, permissions, processes, networking, storage, security, monitoring and troubleshooting — searchable in one place.
190 Terms Searchable Reference
57
September 2026 · #056
Linux Troubleshooting
A runbook, not a tutorial — high load & CPU, memory & OOM, disk & I/O, network issues, and boot/service failures, each with diagnostic commands and root causes.
High Load OOM Disk / I/O
56
September 2026 · #055
Linux Advanced
systemd internals, kernel & sysctl tuning, namespaces & cgroups, performance profiling, and LVM/advanced storage.
systemd Cgroups LVM
55
September 2026 · #054
Linux Fundamentals
Filesystem hierarchy, users & permissions, package management, process basics, and shell/scripting.
Filesystem Permissions Packages
54
September 2026 · #053
Incident Management & Postmortems
Severity classification, Incident Commander role, blameless postmortem writing, and action-item follow-through.
Severity Postmortems SRE
53
August 2026 · #052
CI/CD & GitOps Deep-Dive
Jenkins & GitLab CI, GitOps with ArgoCD/Flux, security gates, canary/blue-green strategy, secrets management.
CI Pipelines GitOps ArgoCD
52
July 2026 · #051
Observability Stack Deep-Dive
Prometheus + Grafana, Loki/ELK, Jaeger/OpenTelemetry, SLO burn-rate alerting, and the OpsCore correlation layer.
Prometheus Loki Jaeger
51
June 2026 · #050
Kubernetes Storage Deep-Dive
PV/PVC, StorageClasses, CSI driver internals, StatefulSets, and a backup & DR strategy that's actually been tested.
Storage CSI StatefulSet
50
May 2026 · #049
Kubernetes & OpenShift Error Runbook
32 error types across 5 layers — Pod, Worker Node, Cluster, API Server, OpenShift — with root causes and production fixes.
Errors Runbook OpenShift
49
April 2026 · #048
Kubernetes Architecture Deep-Dive
Control Plane, Worker Nodes, Storage Nodes, CSI, PersistentVolumes, QoS classes, and scheduling — with interactive diagrams.
Kubernetes Control Plane Architecture
48
March 2026 · #047
Kubernetes Networking Decoded
6 architecture diagrams covering Pod networking, CNI plugins, Ingress, NetworkPolicy, and CoreDNS.
Kubernetes CNI Networking
47
January 2026 · #045
AI/ML Workloads on Kubernetes
GPU node pools, KubeFlow pipelines, model serving with KServe, and GPU resource quotas.
ML Ops Kubernetes GPU
Archive listing · no page yet
45
December 2025 · #044
Platform Engineering & Backstage
Building Internal Developer Platforms with Backstage, service catalogs, scaffolding templates.
Platform Eng Backstage
Archive listing · no page yet
44
November 2025 · #043
GitOps at Scale with ArgoCD & Flux
Multi-cluster deployments, app-of-apps pattern, Flux vs ArgoCD head-to-head comparison.
GitOps ArgoCD Flux
Archive listing · no page yet
43
October 2025 · #042
Observability Stack: Prometheus + Loki + Tempo
The full PLG stack, trace correlation, SLO alerting, and Grafana 11 dashboards from scratch.
Observability Grafana SRE
Archive listing · no page yet
42
Scroll for older issues

THE AUTHOR

One practitioner building in the trenches of Platform Engineering, sharing what actually works.

VA
Vishal Abhinav
· Platform Ops · Kubernetes · Platform Engineering · Cloud Architecture · DevOps · SRE · Infrastructure Automation

I'm a Platform Ops Engineer with 4+ years of experience across Kubernetes, infrastructure, and SRE-engineered systems — designing for failure, automating the boring parts, and writing the runbooks I wish existed the first time something broke at 2 AM. Beyond core K8s and infra, I specialize in notification engine architecture: building and scaling high-throughput messaging systems on SMPP and PDU session handling. This newsletter is the field notes from that work — real runbooks, real postmortems, no theory-only content.

☁️ OCI Certified 🛡️ CISSP Certified 🔶 AWS Certified Cloud Native — In Progress
Specialties
Infrastructure & Automation Specialist SMPP PDU Session
SUBSCRIBE

JOIN 250+
ENGINEERS

One deep-dive every month. No fluff, no spam — just battle-tested knowledge from the Platform Engineering trenches.

Notify me when these land
No spam · Unsubscribe anytime · Published monthly
Prefer a reader? RSS feed