← All casesCase 0214 min

Case 02 · Batch and streaming

The rider moves every second. The accountant counts once a night.

Both are data pipelines. Both are correct. The argument about which one is “better” is the wrong argument, and there is a much more interesting problem hiding underneath it.

Two jobs, two clocks

Both of these are data pipelines.

There is a little bike crawling across a map on somebody’s phone right now. It moves because a rider’s position is being sent onward the moment it is known, and nothing in that chain is allowed to pause and think.

In the same company, a finance team adds up yesterday. Every order, every refund, every commission, every tax line. That job starts at two in the morning, runs for forty minutes, and produces one set of numbers that nobody will question for the rest of the day.

Neither team is doing it wrong. They want different things. The rider’s dot is worthless if it is an hour old. Yesterday’s revenue is worthless if it changes every time someone refreshes.

The argument is never “which is faster”. It is “what is this answer for”.

Most people arriving at this topic expect to learn which one is more modern. That is not the useful part, and it is not what gets asked in interviews either. The useful part is what goes wrong once you pick the fast one.

The first approach

Let it pile up, then deal with it.

Batch means you do not react to anything as it happens. Data lands somewhere and sits there. At an agreed time, a job wakes up, reads a fixed slice of it, does the work, writes the answer, and goes back to sleep.

Below is one Tuesday. Orders come in all day. The batch job ran at 02:00 and will not run again until 02:00 tomorrow. Press play and watch what each side can tell you.

tuesday
00:00before the day starts
batch · runs at 02:00₹0

Waiting for its next run.

answer age —
streaming · always on₹0

Updating as each order lands.

answer age —
Press play. The clock runs one hour per step.

Nothing here is broken. The batch number is not wrong — it is finished. It describes a period that is over and will not change. That is exactly what a finance team wants, and it is why they will not accept a number that moves while they are reading it.

What the batch side cannot do is tell you about right now. By nine at night, its answer is nineteen hours old. If somebody is trying to decide whether to send more riders to Koramangala this evening, that answer is useless to them, no matter how correct it is.

Why anyone still does this

Batch is not the old way. It is the calm way.

It is tempting to read “runs once a night” as outdated. It is worth seeing what you get in exchange.

The input stops moving. The job reads a slice that is over. Yesterday cannot get any more yesterday. Every run on that slice gives the same answer.
Breaking is boring. The job died at 02:14? Start it again at 07:00. Nothing was left half done in a way you have to untangle.
You can go back and redo it. Say someone changes how commission is worked out. You run the job again over the last six months and every number is right. In batch this is a normal day. In streaming it is a project.
You pay only when it runs. Forty minutes of machine time a day, instead of machines kept awake all night doing nothing.

This is why the most trusted tables in most companies are still built by batch jobs — the ones finance signs off, the ones sent to the government. It is not that nobody thought of anything better. It is that “this number will be the same tomorrow” is very hard to get any other way.

The other approach

Never stop. React to each thing as it turns up.

Streaming does the opposite. There is no scheduled run, because the job never stops. Events are produced continuously — a tap in an app, a row changing in a database, a sensor reading, a card being swiped — and something is always listening.

In the middle sits a queue. It holds the events in the order they happened, and it remembers how far each reader has got through them. Your job reads the next event, does its work, and sends the result somewhere a person or another system can see it. Then it reads the next one. It never finishes.

Appsevents happen
→
A queuekeeps the order
→
Your jobalways running
→
Dashboard
or alert
seconds later

Written down like that it sounds like batch with a smaller gap, and that is the mistake almost everyone makes on the way in. It is not a faster batch job. Closing that gap takes away something you did not know you were leaning on.

What you gave up

A batch job knows when the data is complete. A streaming job never does.

This is the whole difficulty, in one sentence, and it is worth reading twice.

When the batch job starts at 02:00, the day it is summarising is over. Everything that was going to happen has happened, and everything that was going to arrive has arrived. The job can count with confidence because there is nothing left to come.

A streaming job has no such moment. At 11:00 it has been handed some of the ten o’clock hour. It has no way of knowing whether that is all of it. Somewhere there is a phone in a lift with an unsent order on it.

Batch waits for the door to close. Streaming has to answer while people are still walking in.

The clock nobody mentions

Every event carries two times, and they are not the same.

When something happened is one time. When your system found out is another. On a good connection they are a second apart and you will never notice. In a basement, a lift, a train tunnel or a crashed app, they are minutes apart — and then everything depends on which of the two you counted by.

Below are nine orders. The solid box is when the order was placed. The dashed box is when it reached your pipeline. The question is simple: how many orders were placed between 10:00 and 11:00?

count by
stop counting at
—you report
—actually happened
—you waited

Solid box = placed. Dashed box = received. The amber line is the moment you stop counting and publish the number.

The uncomfortable bit

There is no cut-off that is both fast and complete.

Go back and try all three cut-offs. You will find there is no setting that wins. Stop at 11:00 and you are quick and wrong. Wait until 11:30 and you are right and late. Every streaming system in the world is sitting somewhere on that line, and somebody chose the spot.

That choice has a name — people call it a watermark, which sounds far more mysterious than it is. It is just the system saying: I will assume nothing older than this is still on its way. It is a guess. A careful guess, and one you are allowed to change.

Guess too early

Late orders miss their hour. Your 10am number is short, and the correction — if there is one — lands after somebody already screenshotted it.

Guess too late

You are right, and too late to matter. The alert you built to catch a problem in seconds now goes off half an hour after anyone needed it.

Guess about right

Which is what real systems do. Plus a rule for the ones that still turn up late: throw them away, or let them quietly fix the number afterwards.

Notice what batch got for free here. When the finance job runs at two in the morning, the phone in the lift has long since reconnected. Batch does not fix the late-data problem. It waits so long that the problem fixes itself.

And the other trap, which is worse. If you count by arrival time instead — the toggle in the widget — the late orders all land in whichever hour they turned up in. Now an order placed at 9:52 is sitting in your ten o’clock total. Nothing crashes. Nothing looks broken. Every hour is a little bit wrong, and it stays that way until somebody puts the dashboard next to the batch report and asks why the two do not match.

The rest of the bill

Three more things that only exist once the job never stops.

Late data is the interesting problem. These next three are the ones that eat a normal working day.

The same event, twice. A network hiccup means a sender retries and the event lands again. In batch you would notice the duplicate rows and clean them. In a stream you have already counted the first one, published it, and moved on. So the job has to remember what it has already seen. Which means it has to remember things at all, and that is where the trouble starts.
Memory that has to survive. A running total, a count per city, “has this card been used in two cities in ten minutes” — all of that is state the job is holding. If the job restarts, all of that has to come back exactly as it was, or your numbers jump. Saving it safely is called checkpointing, and it is most of the reason streaming systems feel heavy.
Fixing the past. Business logic changes. In batch you rerun six months and go home. In streaming, “rerun six months” means replaying six months of events through a live system while it is also handling today, and making sure nothing gets double-counted while you do it. It can be done. It is not an afternoon’s work.

None of this is an argument against streaming. It is an argument against picking it without thinking. Every item on that list becomes somebody’s job at two in the morning, on the night it breaks.

Decide it in one question

What breaks if this answer is an hour old?

Not “how fast can we make it”. Anything can be fast. Ask what actually goes wrong while you wait — that is the thing you are paying for, and streaming is expensive.

Four of those eight want streaming. That ratio is roughly honest for a real company, and it is much lower than the internet implies.

Two questions after that

And then the ones nobody asks until it is too late.

Who is awake when it breaks? A batch job that fails at 02:00 can be rerun at 09:00 by whoever gets in first, and nothing is lost. A streaming job that fails at 02:00 falls further behind every minute, and the pile of unread events keeps growing while somebody looks for the instructions. If the answer to “who is on call” is “nobody, really”, you have your answer about which system to build.

Will you need to redo history? If the rules are new and likely to change, batch will save you months. If everyone already agrees on exactly what you are counting, that argument gets weaker.

These two questions matter more than any speed target, because they are about the team, not the technology. You can swap the tools. You cannot swap the team.

What actually gets built

Almost nobody picks one.

The interview answer people expect is a choice. The real answer, in most companies of any size, is both — with a clear division of labour.

The fast path

For acting

events → queue → stream job
       → alert / live tile

Roughly right, straight away, and allowed to be a bit off. Its job is to make something happen — send a rider, block a card, wake someone up.

The slow path

For agreeing

same events → storage
       → nightly job → warehouse

Complete, repeatable, and the version everyone argues from. If the two do not match, this one is right.

Both read the same events. The difference is when they read them, and what they promise. The live tile might say ₹4.1 lakh at nine in the evening; the next morning’s table says ₹4.06 lakh and that is the number that goes in the deck. Nobody is upset, because everyone was told which one to trust for what.

Fast for deciding. Slow for agreeing.

If you can explain that split in an interview, you are answering a level above the question that was asked.

Same query, two meanings

The code barely changes. The promise does.

Here is an hourly order count. It is the same idea in both worlds.

In batch

Asked once

SELECT hour_of(placed_at) AS hr,
       COUNT(*) AS orders
FROM   orders
WHERE  placed_at >= yesterday
GROUP  BY 1;

Runs, finishes, done. Ask again tomorrow and yesterday’s rows are identical.

In streaming

Never finished

SELECT hour_of(placed_at) AS hr,
       COUNT(*) AS orders
FROM   order_stream
GROUP  BY 1
-- and then: when is an hour done?
-- what about the ones that come late?
-- what if we already published it?

Those three questions in the comments are the actual job. The SQL above them was never the hard part.

Remember this when a tool tells you it lets you “write streaming pipelines in plain SQL”. That is true, and it does help. It does not make the hour finish any sooner.

Three things people say

That sound right and are not.

“Streaming is better, it’s faster.” Faster only counts if something happens during the wait. For a number somebody looks at once each morning over coffee, the extra speed buys nothing — and you still pay for all the extra complexity.
“Batch is the old way.” Salaries, bills, settlements, tax, month-end — huge amounts of money move through batch jobs every night. They do it that way because those numbers have to stop changing. That is not old-fashioned. That is the whole point.
“Streaming means one event at a time.” Plenty of streaming systems quietly gather events into small groups before processing them. Whether the engine does one at a time or two hundred at a time is an implementation detail. What makes it streaming is that it never stops, and it agrees to answer before all the data has arrived.

The third one matters more than it looks, because it means “batch versus streaming” is not a hard border. Run a batch job often enough and it starts to behave like a stream. Make a stream wait long enough and you have built batch again. The real difference is what each one promises about having all the data — not how big the chunks are.

If you get asked

“Design a pipeline for live order tracking.”

The trap is to start naming tools. Tools are the least interesting part of your answer, and the interviewer already knows them all.

Start with the delay. Say what goes wrong during it. “A rider’s position is useless once it is a minute old, so this one has to be a stream.” You have now explained why, instead of just saying so.

Then show you know what it costs. Say that you would bucket by when the event happened rather than when it landed, and that some of it will turn up out of order, or long after the hour it belongs to. That one sentence separates people who have run a streaming pipeline from people who have only read about one.

Then split it. Live path for the map, nightly path for the numbers finance uses. Say plainly that the two will not always match and that this is intended.

Then say what you would not stream. Offering the batch half yourself is the strongest thing you can do, because everyone else waiting outside is arguing for streaming everywhere.

Two questions

Did it land?

Your live dashboard shows 4 orders for the 10am hour. The next morning’s batch table says 7. What happened?

A team wants a real-time dashboard of monthly revenue, refreshed every second. What is the first thing you say?

Your playbook · Case 02

Eight things to keep

  1. Batch waits for the door to closeIt reads a stretch of time that is already over, so its answers stop changing and you can get the same result again.
  2. Streaming answers before the data is completeThat is the real trade. Not speed — the promise it can no longer make.
  3. Every event has two clocksWhen it happened, and when you found out. Count by the first one or your buckets are quietly wrong.
  4. Somebody chose the cut-offA watermark is just a guess about how long to wait for the late ones. Guess early and you miss some. Guess late and nobody needs it any more.
  5. Late data is not the only billRepeat events, memory that must survive a restart, and redoing the past are all your problem now.
  6. Ask what breaks during the delayIf nothing does, you have your answer and it is the cheap one.
  7. Ask who is awake when it failsA stuck stream falls further behind every minute. A failed batch job just gets run again.
  8. Fast for deciding, slow for agreeingMost real systems run both, and say clearly which number everyone argues from.

Your review

How was this case?

Loading ratings…

Next · Case 03

The pipeline that lies to you.

Case 03 is about the job that shows a green tick every morning for three weeks while quietly writing the wrong numbers — and the simple checks that would have caught it on day one.

Open Case 03 →← Case 01