Data readiness for enterprise AI
In the author’s working definition, “AI-ready data” means data an agent or model can use correctly without a human second-guessing every answer. No failure-rate statistics are quoted here: none were sourced.
Needs and where they live
The “In practice” column is the author’s framing (opinion), not sourced; platform-specific statements cite their docs.
| Need | In practice | Layer / vault notes |
|---|---|---|
| Quality and freshness | Tests, SLAs, monitoring; stale or wrong rows produce confident wrong answers | Pipelines and quality tooling; Informatica; lakehouse/warehouse: Databricks, Snowflake |
| Semantics, metadata, definitions | One definition of “revenue”; business terms mapped to columns, metrics and relationships | Semantic models (Power BI), ontologies (knowledge-graph-and-ontology-platforms, ontology), catalogs (Atlan, Collibra, Alation) |
| Lineage and access control | Know where an answer came from; agents inherit the caller’s permissions | Catalogs, Unity Catalog (databricks-genie); Fabric ontology (preview, per Microsoft Learn page dated 2026-10-06) respects OneLake security and source-enforced RLS/OLS/CLS |
| Unstructured content | Parsing, chunking, embeddings, permission-aware retrieval | vector-embeddings, graphrag, retrieval-augmented-generation-overview |
| Evaluation datasets | Question/answer sets and labelled cases to test agents on your data | rag-evaluation, agent-evals-and-observability; AIP Evals (a testing environment for AIP Logic/Chatbot/code functions, Palantir docs) in palantir-foundry-and-aip |
| Cost and token governance | Budgets, usage tracking, caching, model routing per workload | Platform capacity controls; see enterprise-data-ai-platforms-comparison |
| Ownership | Named data owner per domain, review cadence, change process | Organisational; data-product practice; catalog stewardship features |
| Actions and writeback | Agents that change records need governed, auditable actions | palantir-foundry-and-aip action types (Palantir docs: capture operator data or orchestrate decisions that connect to existing systems) |
Checklist (author’s suggestions, opinion)
- Pick one use case and list the exact tables, documents and definitions it needs.
- Make definitions machine-readable (semantic model or ontology) before building the agent.
- Check the agent runs under the user’s identity.
- Build an evaluation set from real questions, rerun on every change.
- Set an owner and freshness expectation per source.
Sources
- Vault notes linked above; Microsoft Learn Fabric ontology (https://learn.microsoft.com/en-us/fabric/iq/ontology/overview, 2026-10-07); Palantir docs (https://www.palantir.com/docs/foundry/ontology/overview/, 2026-10-07); https://www.palantir.com/docs/foundry/aip-evals/overview/ (2026-10-07). Both re-opened 2026-10-07; Fabric ontology confirmed still in preview.
Open items
- No survey or failure-rate figures included; add only with a cited source.
- Tool-by-tool mapping to be extended as other DATA-PLATFORMS notes appear.