Start here
Most posts here explain something complex with one diagram and concise text. I've organised them by borrowing the layers of Clean Architecture — the closer to the core, the more stable and portable (concepts); the further out, the more concrete and replaceable (tools). Dependencies point inward, and so does learning: get the inner concepts solid first, then learn the tools that will eventually be swapped out.
This is the complete map. Posts translated into English are listed by their English title; the rest are marked 中文 and open the Chinese original — 103 of 181 posts are in English so far.
Core concepts: the most stable, most portable layer
This is the layer you keep using through several jobs — the query language, the mental map of data engineering, the principles behind data-intensive systems. Tools get replaced; these don't. Get this solid first, then move outward.
📚 Fundamentals of Data Engineering — Reading Notes · 11 posts
- 1 What Data Engineering Is: Reading Fundamentals of Data Engineering, Ch. 1
- 2 The Data Engineering Lifecycle: Reading Fundamentals of Data Engineering, Ch. 2
- 3 Designing Good Data Architecture: Reading Fundamentals of Data Engineering, Ch. 3
- 4 How to Actually Choose Technology: Reading Fundamentals of Data Engineering, Ch. 4
- 5 Where Data Comes From: Source Systems and Data Generation, Reading Fundamentals of Data Engineering, Ch. 5
- 6 Where to Store Data: The Storage Hierarchy and Its Abstractions, Reading Fundamentals of Data Engineering, Ch. 6
- 7 Moving Data In: Batch or Streaming? Reading Fundamentals of Data Engineering, Ch. 7
- 8 Making Data Useful: Queries, Modeling and Transformation, Reading Fundamentals of Data Engineering, Ch. 8
- 9 The Last Mile of Data: Serving Analytics and ML, Reading Fundamentals of Data Engineering, Ch. 9
- 10 The Chapter That Matters Most and Gets Ignored Most: Security and Privacy, Reading Fundamentals of Data Engineering, Ch. 10
- 11 The Future of Data Engineering: Tools Change, the Foundation Doesn't, Reading Fundamentals of Data Engineering, Ch. 11 (Finale)
📚 SQL: I Thought I Knew It · 12 posts
- 1 Your SQL Doesn't Run in the Order You Wrote It
- 2 The Truth About JOIN: Cartesian Product First, Then Filter
- 3 NULL Isn't a Value, It's "Don't Know"
- 4 GROUP BY: Collapsing Many Rows into One
- 5 Window Functions: Aggregation Without Collapsing
- 6 Deduplication Done Right: DISTINCT Isn't the Only Answer
- 7 Gaps and Islands: Pulling Out Contiguous Ranges
- 8 Time Bucketing and SCD: Two Traps in Handling Time with SQL
- 9 Why Indexes Are Fast — and Why They Stop Working
- 10 Reading EXPLAIN: How the Optimizer Actually Runs Your Query
- 11 Transactions and Isolation Levels: Graded Ways to Keep Concurrency from Fighting
- 12 When SQL Runs on MPP: Greenplum and Cloudberry
📚 Designing Data-Intensive Applications — Reading Notes · 12 posts
- 1 Reliable, Scalable, Maintainable: The Three Goals of a Data System
- 2 Data Models: Relational, Document, Graph — What Are You Actually Choosing?
- 3 Storage Engines: LSM-Trees, B-Trees, and Column-Oriented Storage
- 4 Encoding and Evolution: Letting Old and New Code Read Each Other's Data
- 5 Replication: Single-Leader, Multi-Leader, Leaderless, and Three Replication-Lag Anomalies
- 6 Partitioning: Range or Hash, Where Secondary Indexes Live, and How to Rebalance
- 7 Transactions: The Write Skew Snapshot Isolation Can't Stop, and Three Roads to Serializability
- 8 The Trouble with Distributed Systems: Unreliable Networks, Untrustworthy Clocks, and Half-Dead Nodes
- 9 Consistency and Consensus: Linearizability, the Honest Version of CAP, and Total Order Broadcast
- 10 Batch Processing: The Spirit of MapReduce, Two Roads to a Join, and the Virtue of Immutable Inputs
- 11 Stream Processing: The Dual-Write Trap, CDC, and Stream–Table Duality
- 12 The Future of Data Systems: The Unbundled Database, Kappa, and End-to-End Correctness (Finale)
Transformation and architecture: organising concepts into systems
The layer between concepts and tools — how raw data is refined into usable models, how transformation and layering are arranged, and a few architectural ideas I think are most worth understanding.
Infrastructure: the concrete tools that make the concepts real
The outermost, most concrete, most replaceable layer. Each of these tools turns the concepts above into something that runs — but they are the means, and the core concepts are the end. When looking at a tool, keep asking which concept it implements.
📚 In-memory data structures · Redis — Learning Notes · 12 posts
- 1 What Is Redis? Not Just a Cache, but an In-Memory Data Structure Server
- 2 The Soul of Redis: Five Core Data Structures + Advanced Weapons
- 3 Why Is Single-Threaded Redis So Fast? And the Landmine of O(N) Commands
- 4 Redis Persistence: RDB Snapshots vs AOF Logs, and Whether Data Is Actually Lost
- 5 Redis Expiration and Eviction: TTL, Lazy Deletion and maxmemory Policies
- 6 The Three Cache Disasters: Penetration, Breakdown, Avalanche, and the Right Fixes
- 7 Distributed Locks: From SETNX to Redlock, and That Famous Argument
- 8 Master-Replica Replication: Read/Write Splitting and the Oddities of Replication Lag
- 9 High Availability: How Sentinel Fails Over Automatically
- 10 Redis Cluster: How 16384 Slots Shard and Rescale
- 11 Pipelining, Transactions and Lua: Saving RTTs vs Atomicity
- 12 Pub/Sub vs Stream: Redis's Version of a Messaging System
📚 Event streaming · Kafka — Learning Notes · 5 posts (0 in English)
📚 Distributed computing · Spark — Learning Notes · 6 posts (0 in English)
📚 Workflow orchestration · Airflow — Learning Notes · 9 posts (0 in English)
- 1 Apache Airflow 是什麼?從 cron 到工作流程編排 中文
- 2 跑起第一個 Airflow:Docker 環境 + 你的第一個 DAG 中文
- 3 Airflow 排程的真相:data interval、catchup 與 backfill 中文
- 4 Airflow 任務間怎麼傳資料:XCom、TaskFlow 進階與 params 中文
- 5 Airflow 怎麼連外部系統:Provider、Operator、Hook、Sensor 中文
- 6 Airflow 複雜流程控制:branching、trigger rules、TaskGroup、動態任務 中文
- 7 Airflow 可靠性實戰:冪等、重試、SLA 與告警 中文
- 8 Airflow 測試與部署:別讓一個 typo 弄垮整包 DAG 中文
- 9 Airflow 進階:Datasets、deferrable operators 與 executor 選型 中文
📚 Container orchestration · Kubernetes — Learning Notes · 14 posts
- 1 What Is Kubernetes? From 'Running Containers' to 'Declaring the State You Want'
- 2 Pod, Node, Scheduler: The Three Atoms of a Kubernetes Cluster
- 3 Deployments and Self-Healing: The Reconcile Loop in Practice
- 4 Service: A Fixed Address in Front of Short-Lived Pods
- 5 ConfigMap and Secret: Pulling Configuration and Secrets Out of the Image
- 6 K8s Storage: Volumes, PV/PVC and StatefulSets
- 7 Advanced Scheduling: Getting Pods onto the Right Node
- 8 Airflow + Spark on K8s: How Different Nodes Run Different Pods
- 9 Ingress and Cluster DNS: One Entrance In, One Name to Recognise Each Other
- 10 NetworkPolicy and CNI: The Firewall Between Pods
- 11 RBAC: Who Can Do What to the Cluster
- 12 Cluster Administration: kubeadm, etcd Backups, Upgrades
- 13 Troubleshooting: How to Investigate Pods, Nodes and the Control Plane
- 14 Packaging and Deployment: Helm and Kustomize
📚 Infrastructure as code · Infrastructure as Code — Reading Notes · 8 posts (0 in English)
📚 Configuration management · Ansible for DevOps — Reading Notes · 4 posts (0 in English)
📚 Distributed coordination · standalone
📚 A horizontal view: running data tools as infra (operations / deployment / platform) · Data Tools Through an Infra Lens · 9 posts (0 in English)
- 1 從 Infra 角度看一個工具,要問哪些問題 中文
- 2 Kubernetes:所有東西跑的底座,它自己怎麼站穩 中文
- 3 Kafka:磁碟為王的有狀態叢集 中文
- 4 Redis:記憶體為界的有狀態服務 中文
- 5 RabbitMQ:訊息 broker 的叢集與流控 中文
- 6 Spark:短命 executor 的彈性運算 中文
- 7 Airflow:排程器、worker 與那個藏起來的狀態 中文
- 8 Kafka Connect:連接器的執行時 中文
- 9 把它們兜成一個資料平台 中文
Reliability and observability: across every layer
This layer cuts across everything — data has to be correct, services have to stay up, changes have to be safe, and all of that presupposes being able to see. SRE covers the concepts and culture of reliability (the why); the Grafana LGTM stack covers turning it into real dashboards, alerts and SLOs (the how). It's also where I spend the most effort now, as the DE team's EM and its SRE at once.
📚 Concepts and culture of reliability · Google SRE — Reading Notes · 19 posts
- 1 What Is SRE? Start with the Error Budget
- 2 DevOps vs SRE: One Is an Interface, the Other an Implementation
- 3 SLI / SLO / SLA: A Measurement, a Target, a Contract
- 4 Eliminating Toil: Treat Repetitive Operations as Bugs to Be Killed
- 5 Monitoring: The Four Golden Signals
- 6 Alerting and On-Call: When to Wake Someone Up
- 7 Effective Troubleshooting: Debugging Is a Method, Not a Talent
- 8 Blameless Postmortems: Turning Outages into Organisational Learning
- 9 Incident Response: The Real Enemy in a Major Incident Is Chaos
- 10 Testing for Reliability: Tests Don't Prove the Absence of Bugs, They Let You Move Fast
- 11 Cascading Failures and Overload: Don't Let One Server Take Down the Rest
- 12 Data Pipelines and Data Integrity: Having a Backup Doesn't Mean You Can Restore
- 13 Automation, Release Engineering and Simplicity: Making Change Fast and Safe
- 14 Load Balancing: Pick the Right Datacenter, Then the Right Machine
- 15 Distributed Consensus: How Machines That Crash Agree on One Thing
- 16 An SRE Parachuted into a 'Build Everything In-House' Company: Standing Firm in the First 90 Days
- 17 Reliable cron: The Simplest Scheduled Job Gets Hard the Moment It's Distributed
- 18 Production Readiness Review (PRR): What Makes a Service Worth SRE Taking Over
- 19 Operational Interrupts: What Kills Productivity Isn't the Workload, It's Fragmented Time
📚 Observability tooling: the LGTM stack · Observability with the Grafana LGTM Stack · 8 posts (0 in English)
War stories: rebuilding real systems
Every concept above is tested here against systems I actually built (and actually got wrong). The first battlefield: a live-commerce platform — users ordering by comment, inventory that must never oversell, payments and fulfillment in one chain; each chapter first tells how it was done then, and how I'd design it if starting over.
📚 Re:Building a Live-Commerce Platform from Zero · 21 posts
- 1 The Big Picture: Every System One Comment-Placed Order Touches
- 2 The Opening Move: Five Components and One CI/CD Pipeline
- 3 Comment as Order: Turning a Chat Room into an Order Channel
- 4 Identity and Accounts: Who Exactly Is the Person Commenting?
- 5 Stock: Never Overselling Is This System's One Iron Rule
- 6 From Cart to Order: The State Machine the Hosts Ripped Out
- 7 Third-Party Payments: Three Textbook Pitfalls, and Terrain That Has None
- 8 Pre-Shipment: You Sell a Promise, You Ship a Reality
- 9 The Host and Operations Consoles: There Is No Such Thing as "the Back Office"
- 10 Permissions: Who Can Press Which Button — We Rebuilt It Three Times
- 11 Promotions and Amounts: A Miscalculated Discount Is Harder to Find Than an Oversell
- 12 Notifications: Two Columns, One Scheduled Job, and a Phone Call
- 13 Risk Control and Blocklists: Not an Eviction Notice, a Credit System
- 14 The Moment of Opening: The Three Seconds After the Host Calls a Key
- 15 The Years Without an SRE: A Backend Lead's Production Diary
- 16 Reconciliation: We Never Built It, So Why Did the Books Balance?
- 17 Easy to Upload, Hard to Delete: The Life Cycle of Images and Resources
- 18 Six Engineers Running at the Speed of Twenty
- 19 Microservices for Three People: After Development Was Paused
- 20 A Parallel World: What If It Had Become a SaaS
- 21 Re: If I Really Started Over
People and growth: tech leadership
The other line beyond technology — the lessons of leading people, and, just as important from 2026 on, the lessons of leading AI.
📚 Becoming a Tech Leader — Reading Notes · 8 posts (0 in English)
- 1 領導力 中文
- 2 領導力 - MOI 中文
- 3 領導力 - 成長模型 中文
- 4 領導力 - 成為 Leader 的迷思與痛苦 中文
- 5 領導力 - 創新的三大障礙 中文
- 6 領導力 - 用日記看見自己 中文
- 7 領導力 - 發想力 中文
- 8 領導力 - 願景 中文
📚 The Craft of Working with AI (2026) · 4 posts (0 in English)
Looking for a specific topic? Use the search at the top right, or browse by tag.