Data pipelines for AI (status October 2026)

Four tool classes are often confused. They solve different problems and usually run together.

ClassJobTools in this vaultTypical freshness (illustrative, unsourced)
Batch ELT / ingestionCopy SaaS and database data into the lakehousefivetran, airbyte, matillionminutes to hours
Enterprise integration and qualityIntegration plus quality, MDM, cataloginformatica-idmc, talend-data-fabrichours; CDC where licensed
Streaming and CDCContinuous event transport and stateful processingconfluent-and-kafka (Kafka, Flink)seconds
OrchestrationSchedule, order, retry and observe jobsapache-airflow, dagster, prefect, also kestra, temporaln/a (runs the others)

Rule of thumb (author’s framing, opinion): ingestion tools move data, streaming moves events, orchestrators decide when things run and what depends on what. Transformation (dbt, now one company with Fivetran) sits between ingestion and serving.

What changes for AI

  • Unstructured ingestion. Documents, tickets, chat and files must be parsed, chunked and embedded, not only copied. Some vendors (e.g. Airbyte, Matillion) describe unstructured-source support and agent-context positioning on their own sites; see the tool notes for what was checked.
  • Embedding pipelines. A refresh job re-embeds changed rows or documents into a vector index; orchestrators schedule and gate it (vector-embeddings). Airflow’s Common AI provider (apache-airflow-providers-common-ai, docs show version 0.10.0) runs LLM calls and tool-using agents as tasks.
  • Freshness SLAs. Agents act on current state, so staleness becomes a correctness issue. Stream (Kafka/Flink, Tableflow into Iceberg) or CDC where seconds matter; batch where hours are fine. Dagster expresses freshness as asset policy.
  • Agent-triggered pipelines. Agents can call pipelines through MCP: Prefect Horizon is a platform to deploy, register and govern MCP servers (prefect.io/horizon, no launch date given); MCP servers from Confluent, Airbyte and Qlik, and a Fivetran Context Layer, are vendor-reported but not verified here (see Open items). Governance (RBAC, audit) of those calls is the new control point.
  • Context and lineage. Metadata and lineage (Dagster assets, Informatica CLAIRE metadata) explain which data feeds which model or index.
  • Consolidation (dated). IBM completed buying Confluent (2026-03-17, IBM newsroom; announced 2025-12-08); Fivetran and dbt Labs completed their merger (announced 2025-10-13, completed 2026-06-01); Salesforce completed Informatica (2025-11-18, Salesforce press release); Qlik acquired Talend (completion date not verified); Prefect announced the Dagster acquisition (2026-07-13, Dagster blog; closing not stated). Roadmap effects of these deals are not sourced.

Choosing (opinion)

In the author’s view: start with a managed ingestion tool for SaaS and CDC, add Kafka/Flink only for sub-minute needs, pick one orchestrator, and keep transformation and quality tests in version control. See enterprise-data-ai-platforms-comparison for the platform layer.

Sources

Open items

  • Not verified 2026-10-07: Fivetran Context Layer (private beta, 2026-09-16; fivetran.com press and blog pages did not show it), Qlik-Talend completion date 2023-05-16 (only in the Qlik company note, secondary source), Confluent MCP server and Agent Skills, Airbyte MCP server, Qlik MCP server.
  • Parsing/chunking tools (unstructured ingestion) and vector-database sinks are not covered with sourced detail in this job.