Engineering teams upgrading their analytics infrastructure often frame the decision as batch vs. real-time, with real-time as the obvious destination. The actual decision is more specific: what latency do your use cases require, and what are you willing to pay in infrastructure complexity and cost to achieve it?
Most teams end up with both, and for different tables. Understanding why that is often the right answer requires getting clear on what "real-time" actually means in practice and where batch processing remains the better approach even in 2026.
Defining the latency bands
"Real-time" covers a wide range. A fraud detection model that needs to score a transaction before authorizing it requires sub-100ms latency. A live dashboard showing active user sessions on a website needs latency in the range of seconds to tens of seconds. An operational analytics report that answers "how many orders did we process this hour" can be served with latency of minutes. These are all "real-time" in colloquial usage, but they require completely different architectures.
Sub-second latency requires queries that are pre-computed or served from an in-memory store, not running ad-hoc SQL against a column store in real time. At the seconds-to-tens-of-seconds tier, you can use a streaming processor like Apache Flink or Kafka Streams to maintain continuously updated aggregates that are served directly. At the minutes tier, micro-batch processing (Spark Structured Streaming with a batch interval of 1-5 minutes, or a near-real-time ingestion pipeline that writes to a column store every few minutes) is often sufficient and significantly simpler than a full streaming pipeline.
The question "should we do real-time or batch?" should be asked as: "what is the acceptable latency for each of the specific queries our team runs, and what does each latency tier cost?"
Where batch processing remains the right answer
Batch processing is not legacy. It is the right model for workloads that do not have sub-hour latency requirements, and that category covers a large fraction of analytical queries.
Historical analysis is one example. Queries over the full history of a dataset, running complex joins and aggregations across months or years of data, are well-served by batch processing. The query runs on a schedule (nightly, hourly) against a full snapshot of the data. The compute cluster is provisioned for the job and released when it finishes. The infrastructure cost scales with query complexity and data volume, not with data arrival rate.
Machine learning feature engineering is another. Training a recommendation model on last month's user behavior does not benefit from streaming data. The training pipeline runs on a complete historical snapshot. The complexity of maintaining exactly-once semantics and state management in a streaming feature pipeline is rarely justified by the latency improvement for training workloads.
Complex multi-source joins that combine data from several systems are also often better served by batch. Joining user behavioral data with CRM data with payment history requires that all three sources be available and consistent at query time. Coordinating exactly-once processing and consistent joins across three streaming sources is genuinely hard. Running the join nightly against a stable snapshot of all three sources is simpler, more reliable, and correct for any use case that does not require intra-day freshness.
The specific case for streaming: trigger-driven and freshness-critical use cases
Streaming processing earns its complexity overhead when two conditions hold: the use case requires acting on data within minutes of its arrival, and that action produces measurable value over what batch could provide.
Anomaly detection is the clearest example. A system monitoring ingestion pipeline health to detect when a connector goes silent, or when data volume drops by more than 20 percent, needs to fire an alert within minutes of the anomaly. A batch job running hourly would detect the same anomaly, but the damage from an undetected pipeline failure for an hour is worse than the cost of maintaining the streaming infrastructure.
User-facing dashboards that embed data freshness as a product feature are another case. A product analytics dashboard where users expect to see their last action reflected within 30 seconds is not going to work on hourly batch jobs. The freshness is part of the product value, so the infrastructure cost of streaming is justified by the product requirement, not by a vague sense that real-time is better.
Personalization systems where user actions in the current session should influence what content is surfaced in that same session also require streaming. A recommendation engine that updates user preference signals based on the current browse session and surfaces different content in real time needs sub-minute state updates. Hourly batch updates cannot approximate that.
The hidden costs of streaming that teams underestimate
Streaming pipelines have operational costs that batch pipelines do not. These are not arguments against streaming for the right use cases, but they should enter the decision clearly.
State management is the hardest. Streaming aggregations that maintain running totals or session state need to handle late-arriving data, event reordering, and state expiry. A streaming pipeline that computes daily active users correctly, handles events arriving up to 24 hours late, and produces exactly-once output is significantly more complex to write and debug than a batch SQL query that runs at midnight against the day's complete data.
Exactly-once semantics are non-trivial. "At-least-once" delivery means your streaming pipeline may process the same event multiple times. For use cases where duplication matters (financial calculations, user action counts), you need to either use a framework that provides exactly-once guarantees or design your aggregation functions to be idempotent. Both add complexity.
Infrastructure cost for streaming pipelines is continuous. A batch job runs, consumes compute, and terminates. A Kafka cluster and a Flink application run 24 hours a day. The per-hour cost is lower than a large batch job, but it never stops. For low-volume pipelines where the data arrival rate is modest, the minimum cost of running streaming infrastructure often exceeds the cost of running batch jobs more frequently.
The lambda and kappa architectures as reference points, not prescriptions
Lambda architecture (separate batch and streaming paths that merge at the serving layer) and kappa architecture (single streaming path with long-horizon state serving as the "batch" layer) are useful conceptual references, but neither should be adopted wholesale without evaluating what your specific use cases require.
Lambda's appeal is that the batch layer guarantees correctness for historical queries while the streaming layer provides freshness. The drawback is that you maintain two processing codebases that must produce consistent results. Schema changes must be applied to both. Logic changes must be applied to both. The dual-maintenance burden is real.
Kappa simplifies by having one processing path, but requires that your streaming system can replay historical data efficiently. Kafka's log retention can serve as a replay source, but replaying months or years of data to reconstruct historical aggregates is slow compared to running a batch query against a column store.
For most analytics stacks, the practical answer is neither pure architecture: batch for historical analysis and complex multi-source queries, micro-batch or streaming for freshness-critical metrics, and ad-hoc federation at query time for exploratory analysis that needs to span both. Choosing the architecture by use case rather than applying one model to all queries is not a compromise. It is the correct way to match processing model to latency requirement.
Query across batch and streaming sources together
Nava Labs federates your warehouse, operational databases, and event streams into a single query layer. No ETL required to join across freshness tiers.
Get Early Access