For technology leaders scaling enterprise AI, the biggest constraint often sits beneath the model: data foundations that cannot keep pace with changing business processes and systems. Very pleased to see MIT Technology Review Insights explore this challenge produced in partnership with Uniphore The report points to three areas that shape AI at scale: data readiness, governance and architecture. For enterprise teams, that means: → Automating data discovery and preparation across complex estates → Maintaining control over where data lives and models run → Replacing brittle centralized pipelines with composable infrastructure → Querying intelligence where data already resides → Building an architecture that can evolve as data, models and business processes change Scaling AI requires an architecture built for change from the start. Without a strong data foundation and the right governance, moving from AI pilots to production remains difficult. I share in the report that as data and business processes change, enterprise AI needs an architecture that can evolve with them. With a sovereign architecture, enterprises can build and adapt intelligence on their own terms. Read more here: https://lnkd.in/g8DEW9xb
Scaling Enterprise AI Requires Strong Data Foundations
More Relevant Posts
-
GAPI—GAPI News Feed | Why Data Pruning May Become One of the Most Important Jobs in the AI Era As the global AI race accelerates, most conversations continue to focus on larger models, faster processors, synthetic data generation, and expanding AI infrastructure. Yet one of the most important challenges facing the industry is becoming increasingly clear: Not all data deserves to remain in an AI training ecosystem forever. Just as organizations archive, clean, classify, and retire obsolete records, future AI systems will require continuous data pruning to maintain quality, reduce noise, improve efficiency, and preserve trustworthiness. The next frontier may not simply be creating more data—it may be intelligently deciding what data should be retained, updated, consolidated, versioned, or removed. The Hidden Problem of AI Growth Modern AI systems are consuming unprecedented volumes of information. As models become larger and training pipelines become more sophisticated, organizations face several challenges: * Duplicate content * Contradictory information * Obsolete knowledge * Low-quality synthetic data * Inaccurate metadata * Redundant training records * Unverifiable sources * Version conflicts Without effective pruning, AI systems risk becoming increasingly burdened by data that contributes little value while consuming enormous computational resources. In many ways, the future of AI may resemble maintaining a massive digital ecosystem where growth alone is insufficient. Sustainability requires continuous curation. Why Data Pruning Matters Data pruning can provide several benefits: * Higher training quality * Reduced storage costs * Improved inference performance * Better governance and compliance * More reliable AI outputs * Faster retraining cycles * Reduced energy consumption As AI infrastructure expands globally, pruning may become as important as data acquisition itself. The organizations that can continuously improve the quality of their knowledge assets may gain advantages over those that simply accumulate larger volumes of information. Where GAPI Enters the Picture This is where the vision of the Giga AI Press Initiative (GAPI) becomes particularly interesting. GAPI’s long-term architecture proposes universal orchestration capabilities built upon identity management, metadata structures, registries, digital twins, workflows, and multidimensional mappings. Rather than viewing data as isolated files, GAPI views information as identifiable, traceable, and orchestratable assets. Using concepts such as: * Universal Identity Architecture * Universal Multidimensional Maps (UMM) * Registry Services * Digital Twins * Workflow Orchestration * Metadata Governance * Version Management
To view or add a comment, sign in
-
If data is well governed, it should be ready for AI. I used to assume that almost automatically. I no longer do. Recently, we brought together enterprise documentation, data-platform metadata and code so that an AI system could retrieve knowledge across all three. Much of that information was already governed. Access was controlled, ownership existed, and the sources were known. But as soon as we tried to make it genuinely usable by an AI system, a different set of questions appeared. Which source is authoritative for which type of question? What happens when the documentation says one thing, but the implementation says another? How does the system know whether something is current, deprecated, or only valid in a particular context? Governance wasn't the gap. What was missing was the context that human experts had learned to supply themselves. People know which page is outdated, which source they trust for a particular question, who to ask when two systems disagree, or which implementation actually reflects reality. An AI system doesn't have that tacit knowledge unless we make it explicit: meaning, relationships across systems, freshness, provenance and evidence. And when sources disagree, it also needs to know which one takes precedence, or whether it should answer at all. So the question for data leaders is expanding. Not only "Is this data governed?" but also "Can a machine understand and use it with enough context to produce a trustworthy result?" The closer I get to real enterprise AI use cases, the more the boundaries between data architecture, governance, metadata and AI architecture start to blur. Well-governed data is still the foundation. But AI readiness means making explicit much of what human users previously just knew. What have you had to make explicit for AI that people in your organization used to know implicitly? #DataArchitecture #DataGovernance #EnterpriseAI
To view or add a comment, sign in
-
Spent the day at the AI Enterprise Conference at Pier Sixty. The theme was “From Pilots to Production.” Across sessions, one message came through consistently: Enterprise AI has moved past “Which model should we use?” The real challenge is turning AI into trusted, governed and economically viable production systems. A few things stood out: 1. AI-assisted vs. AI-native - Adding AI to an existing workflow is not the same as redesigning the workflow around AI. - AI-assisted keeps the old process intact and inserts a model into it. - AI-native assumes AI is already part of the operating model. - Most transformation programs are still quietly AI-assisted. 2. The model is no longer the moat. Context is - Most enterprises can access the same frontier models. - Differentiation increasingly comes from data, semantics, business rules, institutional knowledge, retrieval, tool access and lineage. - We are moving from prompt engineering to context engineering. - Which means: the data strategy was always the AI strategy. 3. Agentic AI needs a control plane. - The challenge is no longer simply building agents. - Enterprises need to discover, orchestrate, evaluate, monitor and govern agents running across different platforms and environments. - Without that control plane, agentic AI becomes distributed automation without enterprise control. 4. Human-in-the-loop cannot mean human-on-every-loop. - If humans review every AI decision, the productivity gains disappear into review and rework. - The better model is risk-based autonomy: - Automate predictable, low-risk decisions. - Escalate ambiguity, exceptions and consequential decisions to humans. - As AI absorbs execution, human judgment becomes more valuable, not less. 5. Governance has to run - Governance cannot remain a policy deck. - Identity, approvals, PII controls, evaluations, auditability and policy enforcement need to operate at runtime. A framing I particularly liked: - Governance is a mechanism for scaling human judgment. - Good guardrails do not slow AI down. They create the confidence required to scale it. And there is one more production reality: AI economics matter. - Token usage, context size, agent loops, model routing and inference cost are becoming architecture decisions. - The smartest model for every request is rarely the smartest business decision. My biggest takeaway: - Production AI is becoming an operating-system problem, not a model problem. And these ideas form a dependency chain: - AI-native workflows need a context engine. - Context engines need orchestration and runtime governance. - Governance enables risk-based autonomy. - And autonomy frees human judgment for the decisions that actually require it. That is how we move from AI-assisted enterprises to AI-native enterprises. Which part of that chain is your biggest bottleneck? https://lnkd.in/gDKuKRSp #EnterpriseAI #AgenticAI #AITransformation #AIGovernance#DataStrategy #ContextEngineering
To view or add a comment, sign in
-
Enterprise AI strategy is not just about choosing the right models. It is increasingly about getting the right data, to the right place, in the right quantity, at the right time, with the right context — and under the right controls. A lot of the enterprise AI conversation still starts with models. Which model should we use? Where should we host it? How do we govern it? How do we build agents around it? All important questions. But I increasingly think one of the harder architecture problems is somewhere else: How do we make enterprise data safely and meaningfully consumable by AI? Most enterprises don't have a shortage of data. The challenge is making the right data available to AI in a way that is usable, contextual and governed. And not all data should be made available in the same way. A live account balance may need to come from an authoritative API. Policies and procedures may be better suited to retrieval and RAG. Relationships between customers, products or entities may benefit from a knowledge graph. Real-time changes may need to arrive through events. Curated enterprise datasets may be exposed as governed data products. So the architecture question isn't simply “How do we connect AI to our data?” It is deciding which data should be exposed through which mechanism, for which purpose, with what freshness, context and authority. And as agents move from answering questions to taking actions, another dimension becomes even more important: Just because an AI system can access data doesn't mean it should be allowed to use it for every user, purpose or action. That brings identity, authorization, privacy, lineage, data quality and auditability directly into the AI data architecture. We spent years making enterprise data available to applications and analytics. Now we need to make it safely consumable by AI. That may turn out to be one of the bigger architecture challenges in scaling enterprise AI. #EnterpriseArchitecture #AIArchitecture #DataArchitecture #EnterpriseAI #AIGovernance
To view or add a comment, sign in
-
-
As AI gets deployed more broadly and costs increase, businesses start taking a much closer look at what they’re getting back from their investments. At Snowflake, we call this intelligence efficiency, or how effectively companies turn models, data, context, and compute into business value. As open models continue to improve and drive better price-performance, enterprises have more ways to improve this equation. Committing to a single model provider risks locking them into yesterday's best technology, and missing out on options to improve their economics. To achieve true intelligence efficiency, enterprises need: ❄️ Model choice and flexibility across both frontier and open source providers ❄️ Intelligent routing to ensure tasks are sent to a model that delivers the optimal balance of cost and quality ❄️ Trusted data and context underneath those models so that AI responses are accurate and grounded in a company’s unique information I recently shared with VentureBeat why I believe the economics of enterprise AI depend on optimizing across all three of these areas, and where the industry is headed next. https://lnkd.in/gTkMN_Mg
To view or add a comment, sign in
-
With most orgs having internal AI mandates, companies are only just now seeing the real costs of what company wide usage looks like. Help manage costs and discover new found flexibility with dynamic model routing. ➡️➡️➡️
As AI gets deployed more broadly and costs increase, businesses start taking a much closer look at what they’re getting back from their investments. At Snowflake, we call this intelligence efficiency, or how effectively companies turn models, data, context, and compute into business value. As open models continue to improve and drive better price-performance, enterprises have more ways to improve this equation. Committing to a single model provider risks locking them into yesterday's best technology, and missing out on options to improve their economics. To achieve true intelligence efficiency, enterprises need: ❄️ Model choice and flexibility across both frontier and open source providers ❄️ Intelligent routing to ensure tasks are sent to a model that delivers the optimal balance of cost and quality ❄️ Trusted data and context underneath those models so that AI responses are accurate and grounded in a company’s unique information I recently shared with VentureBeat why I believe the economics of enterprise AI depend on optimizing across all three of these areas, and where the industry is headed next. https://lnkd.in/gTkMN_Mg
To view or add a comment, sign in
-
AI doesn’t need the whole enterprise data landscape to become perfect before it can be useful. But it does need the data to be understandable enough for a model to work with it, and controlled enough for teams to trust what comes out. By now, most teams know that throwing raw, unprepared data into AI is a bad idea. That lesson has had enough production examples. But it is easy to fall into the opposite extreme and assume that every source system has to be cleaned, redesigned, and aligned before AI can move beyond a pilot. While sounding reasonable on paper, it doesn’t match enterprise reality. Data usually comes from platforms built at different times, owned by different teams, and shaped by specific business restrictions. A model can work with that data only if something translates the mess before it reaches the AI layer. I see this through creating a controlled layer between existing sources and downstream AI. This layer reads the data, normalizes the parts that follow a clear pattern, flags what needs review, and applies guardrails before the output reaches systems or users. It also keeps enough context visible so teams can understand what was changed, what was left untouched, and why. This is also why read-only architecture matters in AI readiness work. Diagnosing, mapping, validating, and exporting reviewed improvements is a safer first step than writing changes straight back into production systems. AI Search Readiness Kit applies this pattern in a commerce setting. It prepares catalog data and signals for AI search and assistant scenarios through normalization, harmonization, guardrails, diagnostics, and reviewable outputs, while the existing systems remain in control. Learn more: https://lnkd.in/dEti9jyv For large enterprise AI, perfect data is not a realistic starting point. A controlled path toward interpretable, reviewable, and traceable data usually is.
To view or add a comment, sign in
-
-
🏗️ 𝗔𝗴𝗲𝗻𝘁𝗶𝗰 𝗔𝗜 𝗗𝗼𝗲𝘀 𝗡𝗼𝘁 𝗦𝗰𝗮𝗹𝗲 𝗼𝗻 𝗠𝗼𝗱𝗲𝗹𝘀 𝗔𝗹𝗼𝗻𝗲 | 𝗜𝘁 𝗦𝗰𝗮𝗹𝗲𝘀 𝗼𝗻 𝗙𝗼𝘂𝗻𝗱𝗮𝘁𝗶𝗼𝗻𝘀 ㅤ McKinsey's research on building Agentic AI at scale highlights an important gap between experimentation and enterprise value. ㅤ Nearly two-thirds of organizations surveyed had experimented with AI agents. ㅤ Yet fewer than 10% had scaled an agentic AI use case to the point of delivering tangible business value. ㅤ And 80% reported data-related challenges when scaling AI. ㅤ For me, these numbers point to a broader enterprise architecture issue. ㅤ 𝗔𝗴𝗲𝗻𝘁𝘀 𝗰𝗮𝗻 𝗼𝗻𝗹𝘆 𝘀𝗰𝗮𝗹𝗲 𝗮𝘀 𝘄𝗲𝗹𝗹 𝗮𝘀 𝘁𝗵𝗲 𝗲𝗻𝘁𝗲𝗿𝗽𝗿𝗶𝘀𝗲 𝗳𝗼𝘂𝗻𝗱𝗮𝘁𝗶𝗼𝗻𝘀 𝘀𝘂𝗽𝗽𝗼𝗿𝘁𝗶𝗻𝗴 𝘁𝗵𝗲𝗺. ㅤ Think about the dependency chain: ㅤ 🤖 𝗔𝗴𝗲𝗻𝘁𝘀 ↑ ⚙️ 𝗛𝗶𝗴𝗵-𝗩𝗮𝗹𝘂𝗲 𝗪𝗼𝗿𝗸𝗳𝗹𝗼𝘄𝘀 ↑ 📚 ��𝗿𝘂𝘀𝘁𝗲𝗱 & 𝗔𝗰𝗰𝗲𝘀𝘀𝗶𝗯𝗹𝗲 𝗗𝗮𝘁𝗮 ↑ 🔌 𝗜𝗻𝘁𝗲𝗴𝗿𝗮𝘁𝗶𝗼𝗻 & 𝗘𝗻𝘁𝗲𝗿𝗽𝗿𝗶𝘀𝗲 𝗖𝗮𝗽𝗮𝗯𝗶𝗹𝗶𝘁𝗶𝗲𝘀 ↑ 🏗️ 𝗠𝗼𝗱𝗲𝗿𝗻 𝗔𝗿𝗰𝗵𝗶𝘁𝗲𝗰𝘁𝘂𝗿𝗲 ↑ 🛡️ 𝗚𝗼𝘃𝗲𝗿𝗻𝗮𝗻𝗰𝗲 ↑ 👥 𝗢𝗽𝗲𝗿𝗮𝘁𝗶𝗻𝗴 𝗠𝗼𝗱𝗲𝗹 ㅤ Weakness anywhere in that chain eventually becomes an Agentic AI bottleneck. ㅤ An agent cannot reliably reason over data it cannot trust. ㅤ It cannot execute a business process if the required enterprise capabilities are difficult to access. ㅤ It cannot operate safely without clear permissions, governance and accountability. ㅤ And it cannot become an enterprise capability if nobody owns how it is operated, measured and improved. ㅤ This is why I would be cautious about measuring Agentic AI maturity by the number of agents deployed. ㅤ The better question is: ㅤ 𝗜𝘀 𝘁𝗵𝗲 𝗲𝗻𝘁𝗲𝗿𝗽𝗿𝗶𝘀𝗲 𝗿𝗲𝗮𝗱𝘆 𝘁𝗼 𝘀𝘂𝗽𝗽𝗼𝗿𝘁 𝘁𝗵𝗲𝗺 𝗮𝘁 𝘀𝗰𝗮𝗹𝗲? ㅤ 𝗠𝗬 𝗧𝗔𝗞𝗘𝗔𝗪𝗔𝗬𝗦 ㅤ → The model is only one layer of the enterprise problem. ㅤ → Trusted data, architecture and integration determine how effectively agents can operate at scale. ㅤ → Governance and operating ownership need to evolve alongside technical capability. ㅤ → A successful pilot demonstrates possibility. It does not automatically demonstrate enterprise readiness. ㅤ 𝗗𝗼𝗻'𝘁 𝘀𝗰𝗮𝗹𝗲 𝗮𝗴𝗲𝗻𝘁𝘀 𝗳𝗮𝘀𝘁𝗲𝗿 𝘁𝗵𝗮𝗻 𝘁𝗵𝗲 𝗳𝗼𝘂𝗻𝗱𝗮𝘁𝗶𝗼𝗻𝘀 𝘀𝘂𝗽𝗽𝗼𝗿𝘁𝗶𝗻𝗴 𝘁𝗵𝗲𝗺. ㅤ The technology may be ready. ㅤ The enterprise around it may still have work to do. ㅤㅤ 📖 𝗥𝗘𝗙𝗘𝗥𝗘𝗡𝗖𝗘 ㅤ McKinsey & Company — Building the Foundations for Agentic AI at Scale ㅤ https://lnkd.in/dpDFuUnq ㅤ #AgenticAI #EnterpriseArchitecture #EnterpriseAI #DataStrategy #AIGovernance
To view or add a comment, sign in
-
AI has changed a lot about how we think about data architecture, and one of the biggest changes has been the evolution from data warehouses, to data lakes, and now increasingly to data lakehouses. Data warehouses were built around structured, curated data and were extremely effective for reporting and traditional analytics. But as AI and machine learning became more prevalent, organizations needed access to much larger and more diverse datasets, including semi structured and unstructured data such as documents, images, logs, text, and streaming data. That drove the growth of data lakes. They gave data teams the flexibility and scale to store almost anything and allowed data scientists to work with data that may not have had an obvious business use when it was first collected. But the flexibility of data lakes also created challenges. Data quality, governance, lineage, security, metadata, and performance became increasingly difficult to manage. As AI became more dependent on trusted data, those challenges became even more important. Bad data doesn't just produce a bad report anymore. It can produce a bad model, a bad prediction, or an AI system that provides the wrong answer. This is where the data lakehouse has become so compelling. It brings together the flexibility and scale of a data lake with many of the governance, reliability, and performance capabilities traditionally associated with a data warehouse. AI didn't necessarily cause the move to lakehouses, but it certainly accelerated it. More importantly, AI has changed what we expect from our data platforms. The goal is no longer simply to store data or produce reports. We need a data foundation that can support analytics, machine learning, generative AI, and increasingly autonomous AI applications while maintaining trust in the data. The interesting part is that I don't think the future is about completely replacing warehouses with lakehouses. The lines between the technologies are continuing to blur. The real goal is building a data ecosystem where trusted, governed data can be used wherever the business needs it, whether that's a dashboard, a predictive model, a GenAI application, or an AI agent. #DataEngineering #DataArchitecture #AI #GenerativeAI #DataLakehouse #MachineLearning #DataStrategy
To view or add a comment, sign in
-
Most enterprise AI projects hit a wall not because of big data, but because of small data. Everyone talks about governing massive data lakes. Terabytes of customer transactions. Petabytes of sensor readings. But the real headaches? They often come from the tiny, critical datasets. Think about it: 1. Feature Stores: These are often small, highly curated datasets. Yet, a single inconsistency or bias in a feature can ripple through every model that uses it. Governance here isn't about volume, it's about precision and lineage. 2. Training Labels: Who created them? What were the instructions? Were they reviewed? Small, inconsistent labeling datasets introduce silent biases that are incredibly hard to debug later. Your AI learns from these small samples. 3. Metadata: This is the DNA of your data. Data types, definitions, ownership, access rules. When metadata is messy or incomplete, even the most robust big data governance strategy falls apart. It's the context for everything else. Focusing only on the scale of data misses the point. The influence isn't always proportional to volume. For enterprise AI, the quality and consistency of these 'small data' elements are often the true determinants of success or failure. They're where subtle biases creep in, where explainability breaks down, and where trust in AI erodes. We need to treat these critical small datasets with the same, if not more, rigor than our largest data assets.
To view or add a comment, sign in
Explore related topics
- How to Scale AI in Enterprises
- Developing A Governance Structure For AI
- How To Scale AI In Regulated Industries
- Building Scalable AI Infrastructure
- Choosing The Right AI Models For Enterprises
- Scaling AI While Maintaining Compliance
- Challenges of Scaling Artificial Intelligence
- The Role Of Leadership In AI Scaling
- How to Scale Foundation Models for AI Infrastructure
- Ensuring Data Quality For Scalable AI