Data readiness for enterprise AI

In the author’s working definition, “AI-ready data” means data an agent or model can use correctly without a human second-guessing every answer. No failure-rate statistics are quoted here: none were sourced.

Needs and where they live

The “In practice” column is the author’s framing (opinion), not sourced; platform-specific statements cite their docs.

NeedIn practiceLayer / vault notes
Quality and freshnessTests, SLAs, monitoring; stale or wrong rows produce confident wrong answersPipelines and quality tooling; Informatica; lakehouse/warehouse: Databricks, Snowflake
Semantics, metadata, definitionsOne definition of “revenue”; business terms mapped to columns, metrics and relationshipsSemantic models (Power BI), ontologies (knowledge-graph-and-ontology-platforms, ontology), catalogs (Atlan, Collibra, Alation)
Lineage and access controlKnow where an answer came from; agents inherit the caller’s permissionsCatalogs, Unity Catalog (databricks-genie); Fabric ontology (preview, per Microsoft Learn page dated 2026-10-06) respects OneLake security and source-enforced RLS/OLS/CLS
Unstructured contentParsing, chunking, embeddings, permission-aware retrievalvector-embeddings, graphrag, retrieval-augmented-generation-overview
Evaluation datasetsQuestion/answer sets and labelled cases to test agents on your datarag-evaluation, agent-evals-and-observability; AIP Evals (a testing environment for AIP Logic/Chatbot/code functions, Palantir docs) in palantir-foundry-and-aip
Cost and token governanceBudgets, usage tracking, caching, model routing per workloadPlatform capacity controls; see enterprise-data-ai-platforms-comparison
OwnershipNamed data owner per domain, review cadence, change processOrganisational; data-product practice; catalog stewardship features
Actions and writebackAgents that change records need governed, auditable actionspalantir-foundry-and-aip action types (Palantir docs: capture operator data or orchestrate decisions that connect to existing systems)

Checklist (author’s suggestions, opinion)

  1. Pick one use case and list the exact tables, documents and definitions it needs.
  2. Make definitions machine-readable (semantic model or ontology) before building the agent.
  3. Check the agent runs under the user’s identity.
  4. Build an evaluation set from real questions, rerun on every change.
  5. Set an owner and freshness expectation per source.

Sources

Open items

  • No survey or failure-rate figures included; add only with a cited source.
  • Tool-by-tool mapping to be extended as other DATA-PLATFORMS notes appear.