For thirty years, integrating data has meant the same thing: copying it from one place to another. ETL extracts it and reloads it into a warehouse, the ESB routes it through a central hub, the data lake dumps it all in one spot « just in case. » Three generations of tools, one single reflex: capture, duplicate, store. But in 2026, data has become real-time, distributed and regulated. And the old premise of « storing in order to integrate » is now paid for in latency, hidden costs and attack surface.
ETL, ESB, Data Lake: three answers, one shared premise
The three historical pillars of integration address the same problem at different points in time. ETL (Extract, Transform, Load) was born with data warehouses: it extracts data from source systems, transforms it, then reloads it into a central database, most often overnight, in batches. The ESB (Enterprise Service Bus) generalized the integration hub: a central bus through which everything transits, with its message queues and intermediate copies. The data lake, finally, pushed the logic to the extreme: dump all raw data into a single reservoir, betting that we will know how to exploit it later.
The common thread is obvious: in all three cases, to move data around, you start by copying it and putting it somewhere. This market remains massive and growing. Data integration was worth 17.58 billion dollars in 2025 and is expected to reach 33.24 billion in 2030, a yearly growth rate of 13.6% (source: MarketsandMarkets). But within this market, the center of gravity is shifting.
The segment that is exploding is cloud and real-time integration. The iPaaS (Integration Platform as a Service) market exceeded 9 billion dollars in revenue in 2024, up from 7.8 billion in 2023 and 5.9 billion in 2022, and Gartner sees it crossing 17 billion by 2028 (source: Gartner, via Informatica). For comparison, the entire application middleware market grew 11.9% in 2024 to reach 64.1 billion dollars (source: Gartner): modern integration is growing roughly twice as fast as classic middleware. The direction is clear, even where the old tools remain in place.
Batch can no longer keep up: data has become real-time
The ETL model rests on an assumption that has become false: that data can wait. Yet it no longer waits. IDC estimated that the global volume of data would grow from 33 zettabytes in 2018 to 175 zettabytes in 2025, and above all that close to 30% of that data would need to be consumed in real time (source: IDC, Data Age 2025). Data that is decided by the second sits poorly with overnight batch processing.
This data is not only faster, it is also elsewhere. As early as 2018, Gartner predicted that 75% of enterprise data would be created and processed outside a centralized data center or a traditional cloud by 2025, compared with around 10% at the time (source: Gartner). Mechanically pulling all of this data back to a central warehouse to process it becomes a geographic as much as an economic absurdity.
The market confirms this shift. Streaming analytics, meaning the processing of data while it is in motion, is expected to grow from 29.53 billion dollars in 2024 to 125.85 billion in 2029, a yearly growth rate of 33.6% (source: MarketsandMarkets). On the ground, more than 80% of Fortune 100 companies already use Apache Kafka to move their event streams (source: Apache Kafka). And the pressure keeps rising: Gartner predicts that the adoption of data streaming for agentic AI will exceed 60% by 2028, compared with less than 15% in 2025 (source: Gartner).
The gap between perceived value and reality nonetheless remains enormous. A survey conducted for Solace showed that 85% of organizations recognize the business value of event-driven architecture, but that only 13% believe they have reached its full maturity (source: Coleman Parkes for Solace). Real time is understood and wanted. What is missing is a way to reach it without building yet another warehouse.
The data lake, from reservoir to swamp
The data lake promised to solve the problem by removing silos: a single place to store everything, and we will sort it out later. Yet Gartner warned as early as 2014 with its « data lake fallacy »: without descriptive metadata or governance, a data lake turns into a « data swamp, » a marsh where information exists but becomes impossible to find and impossible to use (source: Gartner).
Ten years later, the diagnosis is borne out in the numbers. Stored data remains very largely dormant.
- 55% of an organization’s data is « dark data »: it exists but is neither found, nor prepared, nor analyzed (Splunk, The State of Dark Data)
- 68% of the data available to companies is never used, according to a survey of 1,500 executives (Seagate and IDC, Rethink Data)
- 60 to 73% of all of a company’s data goes unused for analytics (Forrester)
The problem is not volume, it is the governance of everything that lies dormant. Gartner in fact predicts that 80% of data and analytics governance initiatives will fail by 2027 (source: Gartner). Accumulating has never been the same as exploiting. The larger the reservoir grows, the more it costs to maintain, secure and govern, for a fraction of the value actually used.
What moving data really costs
Copying and storing is not neutral, neither financially nor in terms of risk. Every copy that leaves a cloud is billed: AWS standard outbound transfer starts at 0.09 dollar per gigabyte beyond the free 100 GB per month (source: AWS). Trivial per unit, this cost becomes structural at scale: Gartner estimates that most customers devote 10 to 15% of their cloud bill to egress fees alone (source: Gartner, via Fierce Network). On worldwide public cloud spending forecast at 723.4 billion dollars in 2025 (source: Gartner), the total is staggering. Lawmakers have taken up the issue: the European Data Act will fully remove switching fees, including egress fees, as of January 12, 2027 (source: European Commission).
The most dangerous hidden cost, however, is not on the bill, it is in the risk. Every copy is one more target. IBM’s report on the cost of data breaches indicates that 35% of leaks involve « shadow data, » this data duplicated outside the security perimeter, and that these breaches cost on average 16% more (source: IBM, Cost of a Data Breach 2024). Knowing that the global average cost of a breach reached a peak of 4.88 million dollars in 2024 (source: IBM), every needless replica becomes a security debt.
Then there is sovereignty. The three American giants concentrate 70% of the European cloud market, while the share of European providers has fallen to around 15% (source: Synergy Research Group). Yet many integration and iPaaS tools are American services, subject to the Cloud Act. In their joint legal assessment, the EDPB and the EDPS conclude that, absent an international agreement, a provider subject to Union law cannot legally transfer data to the United States on the basis of such requests, in direct contradiction with Article 48 of the GDPR (source: EDPB and EDPS). Routing your data through an American integration hub means exposing every copied byte. As we explained regarding digital sovereignty, the provider’s nationality takes precedence over the location of the servers.
In-transit orchestration: integrate without storing
If copying and storing is the source of the problem, the solution is to stop doing it. That is the principle of in-transit orchestration: processing data while it is in motion, at the precise moment it passes through, without ever putting it down or duplicating it. Data is no longer pulled toward a central reservoir, it is transformed on the fly then released to its destination.
This is the approach of iD4Connect. Each DataCell is an autonomous processing unit that cleans, transforms, enriches, anonymizes or routes data during its transit, then keeps nothing. The orchestration of these DataCells, organized into a DataGraph, makes it possible to build complex flows between any applications, on-premise, in the cloud or in hybrid setups. The Universal Connector handles ingestion via all standard protocols (REST, SQL, MQTT, Kafka, OPC-UA, SFTP…), without imposing its own format.
The benefit is threefold. The exposure surface drops to zero, since no data is stored along the way. Sovereignty is guaranteed by design, the architecture being conceived in Europe and executed within the customer’s perimeter. And the cost becomes predictable, with no warehouse to maintain and no egress to pay for copies that no one uses. Where ETL, the ESB and the data lake add a storage layer between your systems, in-transit orchestration removes it. To go further, we detail this how we compare to existing tools.
ETL, the ESB and the data lake did not fail: they solved the problems of their era. But that era assumed slow, centralized and lightly regulated data. In 2026, data is fast, everywhere and under legal scrutiny. The real question is no longer « where to store in order to integrate, » but « how to integrate without storing. » And the best copy remains the one that was never made.
Discover how iD4Connect orchestrates your data without intermediate storage →