Does linking knowledge help AI navigate

Do Document Links Help AI Find What It Needs?

When our team started running an AI agent, one question came up: How do we give it accurate context about our systems? We put Markdown documents in a Git repository and linked related documents together. No vector database, embeddings, or separate search system. The main purpose was to give the AI agent something to read and refer to, but we also wanted the knowledge to be easy for people to find and maintain. ...

September 12, 2026 · 7 min · Jaehyuk Jang
Debezium CDC silent data errors

Debezium Does Not Replicate the Source As-Is: Three Ways Data Can Drift Quietly

In the company tech blog, I previously wrote about redesigning our CDC pipeline with Debezium and Flink, and later about moving our replication model toward CDC-based incremental processing. The structure was simple: Debezium captured changes from MySQL, Kafka carried the events, and Flink replicated them into S3 and Iceberg tables. I will skip the architecture and the reasons behind that design here. This post is about what happened after that. The issues we ran into while operating Debezium fell roughly into two groups. One was data quietly drifting from what we expected. The other was connector behavior that did not match our operational assumptions. In this post, I will focus on the first group: cases where the data itself ended up different. ...

September 5, 2026 · 10 min · Jaehyuk Jang
CDC Small File Problem

The Hidden Cost of CDC Pipelines: How Small Files Create an S3 Request Bomb

In Why We Redesigned Our CDC Pipeline with Debezium and Flink, we introduced our Debezium + Flink-based CDC pipeline. Debezium captures database changes, routes them through Kafka, and Flink writes Parquet files to S3 every 5 minutes. The pipeline worked well for over a year. Data consistency was verified, and operations were stable. The problem remained invisible until we dug into our S3 costs as part of a data platform TCO analysis. ...

June 11, 2026 · 8 min · Jaehyuk Jang