yudopr.dev
Back to all posts

The Data Stack I Designed vs. The One I Inherited

2026-09-29•11 min read
Data EngineeringData WarehouseCDCSchema DriftClickHouseCareer

The Data Stack I Designed vs. The One I Inherited

Two months ago I joined a new company, also a fintech, as a data engineer. Same industry, same job title on paper, and a completely different planet underneath. I spent the first few weeks in a low-grade state of shock, and it took me a while to admit what was actually bothering me.

The realization is embarrassingly simple: I had quietly forgotten how wide the data-tech spectrum is, and that not every company runs the "proper" stack I was used to designing. My previous company had a mature, well-funded warehouse that I'd helped shape as the lead. This one runs on a lot of manual Python scripts. Both are "fintech data platforms." The gap between them is enormous, and I'd been living so deep inside one of them that I'd stopped believing the other kind exists.

This is a blunt, honest retelling of that gap. No employer names, no colleague names, no blame — just what I walked into, what it taught me, and the parts of my own thinking that turned out to be narrower than I assumed.

The stack I was used to designing

In my previous role I was a Data Engineer Lead. I designed the data warehouse system end to end, and I optimized for one thing above all: a new engineer or an intern could open it and understand it without me narrating.

That constraint pushed the whole architecture. We had schemas, documented pipelines, orchestrated jobs, and a UI you could click into for most steps, because a pipeline you can only inspect by reading code is a pipeline only its author can debug. I relied on a stack that was open source but had a real interface — a dashboard, a job graph, a schema browser. Not because open source is automatically better, but because the absence of a UI turns every incident into archaeology.

One more real constraint shaped that stack, and I don't want to gloss over it: the company would not spend on licenses. The budget line for software was effectively closed; the money went to VMs and cloud. So the stack I designed leaned hard on open-source tooling with usable interfaces, running on machines we controlled. Debezium and Kafka Connect for CDC, Kafka as the backbone, Parquet on object storage for the lake, with zstd or lz4 compression, and compressed columnar data because storage math is storage math no matter what tool writes it.

I didn't resent the constraint. In some ways it made me sharper about what actually matters versus what just looks expensive. But here's the part I got wrong, in hindsight: I let a set of very specific constraints harden into a belief about how data platforms "should" be built. I thought my stack wasn't a choice, it was the correct answer that happened to be cheap. I stopped being able to see a company that solved the same problems completely differently and just fine.

The gap, and then the shock

I left that role and rested for about two months. Genuinely rested — not a grind between interviews. It was the right call, and I'd recommend it.

Then I came back to data, at another fintech. I thought I knew what to expect: some version of the patterns I'd been designing, maybe rougher around the edges. I was not prepared for manual.

Not manual in the sense of "a few cron jobs." Manual in the sense that the core of the ingestion layer was hand-rolled Python scripts that someone ran or scheduled, one table at a time, with logic living in the script. Let me be specific about the contrast, because this is the whole post in miniature:

  • Where I came from: CDC via Debezium + Kafka Connect feeding Kafka. Schema evolution was a property of the system, not something I had to maintain.
  • Where I landed: CDC as Python scripts. Every one of them was a place where schema evolution was now my problem, personally, by hand.

That single difference is the gap. Everything else in this post flows out of it.

Schema evolution: the feature I didn't know I was buying

The lesson that really crystallized for me was schema evolution, because I hit it the hard way, and recently.

I needed to ingest a new table, so I wrote the script. Yesterday. Standard shape: read the source, map the columns, write them into the warehouse. It wasn't deployed yet — still in review. So the next day, before it went out, I went back to double-check the source, and the DDL at the source had changed. A column had been altered. In twenty-four hours, on a table that hadn't even reached production ingestion yet.

If I'd deployed that script as written, it would have broken on data that didn't match the shape it expected. And here's the nasty part: in the Debezium setup I came from, this would not have been a problem. Debezium tracks the source schema, Kafka Connect propagates the change through the topic, and the downstream schema history handles it. The system absorbs the change. I never would have noticed it, because noticing drift was the tool's job, not mine.

In the Python-script world, there is no schema history. If the source changes, my script is now wrong, and the only way I find out is when it fails in production, or when a human notices. The system doesn't protect the pipeline; the pipeline has to protect itself. That's a fundamentally different engineering relationship with your data.

The audit: schema drift everywhere

That one bad experience made me curious, and curiosity turned into an audit of the current warehouse's overall state. Which is, bluntly, not good.

There's a lot of schema drift — the warehouse's shape has quietly diverged from the source in many places, for many tables, over time. Some of it is because scripts were written against a schema that later changed and nobody revisited them. Some of it is because there was never a single authoritative definition of the expected shape in the first place, so every script is its own little island of truth that slowly goes stale.

Standing in the middle of that mess, I had a very humbling thought. A large part of what I called "my expertise" at the last job was not model-building or SQL craft — it was having a system that made schema drift visible and survivable by default. Debezium, Kafka, schema history, an orchestrator with a UI: together, those tools are a drift-detection machine. Strip them out, replace them with scripts, and that protection is gone, and you feel the absence immediately as a junior-ish feeling of "how did anyone ever run this without a safety net?" It's not glamorous expertise, but it's the kind you only notice you're missing when you don't have it.

Storage: the lake I knew vs. the warehouse I got

Storage was the other place the two worlds diverged, and it made the "spectrum, not ladder" point concrete for me.

Where I came from: a data lake on object storage, with data written as Parquet — highly compressed, columnar, with zstd or lz4 encoding on the cloud storage itself. Compressed, cheap, and the format is excellent at the thing lakes are for: storing a lot of history for not much money.

Where I landed: ClickHouse as the warehouse. And I want to be fair here, because ClickHouse is genuinely good. It's fast, it's columnar, it compresses well on its own, and for interactive analytical queries it's excellent. I'm not here to trash a good tool.

My issue isn't the tool — it's the lifecycle. Everything is kept, nothing expires, and nothing is ever moved down to cold storage. So the warehouse accumulates all history forever in the hot, expensive tier, when most of it is old enough that nobody queries it. A lake handles this naturally — the Parquet sits in cold, cheap object storage, and you pay for it like cold storage should be paid for. Here, cold data is being stored at hot prices forever, and the bill quietly compounds. It's not a correctness bug; it's a design that ignores the tiering that makes long-term data affordable.

The real lesson: a spectrum, not a ladder

This is the part I actually came away with, and it's more humbling than I expected.

I used to think of data stacks as a ladder — a progression where Debezium + Kafka + a compressed lake sits above Python scripts, and any reasonable company would eventually climb toward the good stuff. Joining this company reset that. There is no single correct data stack. There's a wide spectrum, and most companies sit wherever their constraints, money, history, and people put them. Both my old stack and this new one are valid answers to their own constraints. Neither is universally right.

The specific failure I made wasn't designing a good stack under a budget constraint. That part was fine. The failure was mistaking my locally-optimal, constraint-shaped stack for the general answer — and then being genuinely surprised when a real company didn't match it. The lesson isn't "I should have known better about other architectures." It's more like: your best practice is local. It is the best practice for the constraints you're under, and the moment you carry it somewhere else, those constraints don't travel with you and neither should the certainty.

If I'd understood that two years ago, I probably would've walked into this role calmer, and a lot more curious instead of shocked.

What I'd actually fix first

If I were prioritizing the current mess with a realistic budget, I would deliberately not start by replacing ClickHouse or rewriting everything in Kafka. That's the glamorous move and the wrong first move. Unglamorous and correct comes first:

  • Get a schema contract in front of every ingestion. Each table's expected shape defined, versioned, and checked in one place, so drift surfaces before it breaks a pipeline instead of after.
  • Make the manual scripts defensive. They don't have Kafka's schema history, so they need explicit shape validation and loud failure instead of silent wrongness.
  • Add retention and tiering. Move old data to cheap cold storage and let the hot tier hold what people actually query. Fix the economics before adding more machinery.
  • Then, and only then, consider upgrading the pipeline substrate — e.g. moving hand-rolled CDC onto something that handles schema evolution natively — once the cheap wins are banked.

Fix the drift and the lifecycle first. The platform rewrite is a later chapter, not the first one.

Wrapping up

I'm not writing this to make my old company look good or the new one look bad. Both are just data stacks that exist in the real, wide spectrum of how data actually gets done.

What I did lose, temporarily, was perspective. I'd been the person who designed the nice stack for so long that I'd stopped being able to see a company that runs it the other way. Getting handed the messy one didn't make me a worse engineer; it just reminded me that the elegant stack I was proud of was never the answer. It was an answer, shaped by a budget, a team, and a set of constraints that don't apply everywhere.

So if you're early in your data career and your company runs something that looks less sophisticated than what you've read about: that's normal, that's a spectrum, and it's not a referendum on your skills. And if you're senior and feeling smug about your stack — I was too. The honest move is to hold your architecture as one good answer among several, and stay curious about the ones you haven't seen. The gap between the stack I designed and the one I inherited wasn't a failure of anyone's standards. It was just the reminder that data engineering is a lot wider than the corner I was standing in.