Event Streams, Log Compaction and Derived Systems Explained
A database can tell you the current state of your data.
An event stream can tell you how that state was created.
That distinction is simple, but it becomes important when designing event-driven and distributed systems.
Consider an inventory example.
A product starts with some stock and receives these changes:
+10 added → -2 sold → +5 added
The final stock is:
13
A database might simply store:
Product A → Stock = 13
That is useful for an application serving the current inventory.
But the database value alone does not describe the sequence of changes that produced it.
An event stream can preserve those changes.
This leads to three related concepts:
- Event streams preserve changes over time.
- Log compaction can retain the latest value for each key.
- Derived systems transform source data into views optimized for particular applications.
Understanding how these concepts fit together is useful when designing modern data and backend architectures.
Event Stream vs. Database
The simplest distinction is:
Database = current state
Event stream = sequence of changes
Suppose an inventory system receives these events:
+10 units -2 units +5 units
The resulting state is:
Stock = 13
A database is generally interested in serving that current state efficiently.
An event stream preserves the events that produced it.
The two representations answer different questions.
Database
What is the stock now?
Event stream
What changes happened to produce that stock?
Neither representation is automatically better.
They serve different purposes.
Why Keep an Event Stream?
An event stream becomes useful when the history of changes has value.
For example, historical events can support:
- Auditing
- Debugging
- Recovery
- Replay
- Event-driven processing
- Multiple downstream consumers
- Rebuilding derived views
Imagine a product update needs to reach several systems.
The same event could potentially be consumed by:
Inventory Service ↓ Search Index ↓ Analytics System ↓ Reporting
Instead of tightly coupling every system to one database implementation, the event stream can act as a source of changes for downstream processing.
This can provide architectural flexibility.
However, it also introduces additional infrastructure and operational responsibility.
Event Streaming Is Not Always Necessary
One common mistake is assuming that every application should use event streaming.
A simple CRUD application may only need:
Application → Database
Adding an event broker and asynchronous consumers could introduce unnecessary complexity.
The team may then need to manage:
- Event schemas
- Consumers
- Retention
- Ordering
- Replay
- Monitoring
- Failure handling
- Duplicate processing
For a small application, those costs may not be justified.
The better question is:
What problem does the event stream solve?
If the application needs historical changes, asynchronous processing, multiple consumers, or rebuildable derived systems, the additional complexity may be worthwhile.
What Is Log Compaction?
Now consider a system that does not need every historical update.
Suppose an e-commerce search system only needs the latest product information.
The event history might contain:
Product A → +10 Product A → -2 Product A → +5
The current state is:
Product A → Stock = 13
If the downstream system only needs the latest value, processing the complete history every time may be unnecessary.
This is where log compaction becomes useful.
Conceptually, a compacted log can retain the latest value for a key while allowing older superseded values to be removed over time.
Instead of requiring the downstream system to process every historical update, it can work with the latest available state.
Why Log Compaction Helps Derived Systems
Consider a search index.
The search system may need:
- Product name
- Product price
- Product category
- Current stock
It may not need every historical product update.
A compacted representation can make rebuilding the latest state more efficient.
This is especially useful when a downstream system needs to reconstruct its state after:
- Failure
- Restart
- Data loss
- Reprocessing
- Deployment changes
The important distinction is:
Full event history is useful when history matters.
Compacted state is useful when the latest value matters.
What Are Derived Systems?
A derived system is a system or data representation created from another source of data.
Examples include:
- Search indexes
- Materialized views
- Analytics tables
- Product catalog views
- Application-specific read models
Consider this flow:
Event Stream ↓ Stream Processing ↓ Derived State ↓ Search / Application
The derived system is optimized for its specific workload.
A search index, for example, is designed for efficient searching.
An analytics table may be optimized for reporting queries.
A materialized view may be optimized for a particular read pattern.
The derived system does not necessarily need to become the authoritative source of every historical change.
Event Stream + Compaction + Derived System
These concepts become easier to understand when viewed together.
A simplified architecture might look like this:
Event Stream │ ┌─────────┴─────────┐ │ │ Full History Compacted State │ │ │ ↓ │ Derived System │ │ ↓ ↓ Audit / Replay Search / Queries
The event stream provides the change history.
The compacted representation provides the latest state per key.
The derived system consumes that information and creates a representation optimized for a particular use case.
A Practical Example: Product Inventory
Consider an e-commerce platform.
A product starts with:
Stock = 0
Then several events occur:
+10 units received -2 units sold +5 units received -1 unit sold
The resulting state is:
Stock = 12
Different parts of the platform may need different information.
Inventory Service
Needs:
Current stock = 12
Audit System
Needs:
The complete sequence of inventory changes
Search Index
May only need:
Product A → Stock = 12
Analytics System
May need:
Aggregated inventory events by product, category, or time period
One underlying stream of changes can therefore support multiple representations.
The Architectural Trade-Off
This flexibility is useful, but it is not free.
An event-driven architecture can introduce new concerns.
Event Ordering
Consumers may need to process events in the correct order for particular keys.
Duplicate Processing
Consumers may process the same event more than once, so important operations often need careful idempotency design.
Schema Evolution
Events can live for a long time, which means changes to event formats need to be managed carefully.
Replay
If events are used to rebuild derived systems, the system needs a reliable replay strategy.
Retention
Teams need to decide how long historical events should remain available.
Compaction
Compaction changes what historical information remains readily available in the compacted representation.
These are architecture decisions, not simply configuration details.
Full History vs. Latest State
A useful mental model is:
| Requirement | Useful Representation |
|---|---|
| Current state | Database / state store |
| Complete change history | Event stream |
| Latest value per key | Compacted log |
| Search-optimized data | Search index |
| Analytics | Derived analytics tables |
| Application-specific reads | Materialized view / read model |
The same business data may therefore have multiple representations.
That is not necessarily duplication for its own sake.
Each representation can be optimized for a different purpose.
When Should You Use This Architecture?
Consider event streaming when you have requirements such as:
- Multiple independent consumers
- Historical event processing
- Asynchronous workflows
- Replay or recovery requirements
- Derived read models
- Real-time stream processing
- Audit requirements
A simpler database-centric design may be more appropriate when:
- The application is primarily CRUD
- Historical changes are not important
- There are few downstream consumers
- Operational simplicity is the priority
- The workload does not require asynchronous event processing
The goal should not be to maximize architectural sophistication.
The goal should be to solve the actual system requirements.
Key Takeaways
- Databases are commonly used to represent current application state.
- Event streams preserve sequences of changes.
- Log compaction can maintain the latest value for each key.
- Derived systems transform source data into representations optimized for specific workloads.
- Event streams can support replay, auditing, recovery, and decoupled consumers.
- Not every application needs event-driven architecture.
- More infrastructure brings more operational responsibility.
- The right architecture depends on the problem being solved.
Final Thought
A useful way to think about modern data architectures is to separate:
What happened?
from:
What is true now?
And sometimes add a third question:
How should this data be represented for this particular application?
Event streams help answer the first.
Databases and compacted state help answer the second.
Derived systems help answer the third.
The real architectural skill is knowing when these representations should exist—and when they don't need to.
What would you choose for a system that needs both current state and historical changes:
Database only, event stream, or both?
