𝗜𝗳 𝘆𝗼𝘂 𝘄𝗮𝗻𝘁 𝘁𝗼 𝗯𝘂𝗶𝗹𝗱 𝗮𝗻 𝗔𝗜 𝘀𝘁𝗿𝗮𝘁𝗲𝗴𝘆 𝗳𝗼𝗿 𝘆𝗼𝘂𝗿 𝗰𝗼𝗺𝗽𝗮𝗻𝘆, 𝘆𝗼𝘂 𝗳𝗶𝗿𝘀𝘁 𝗻𝗲𝗲𝗱 𝘁𝗼 𝗯𝘂𝗶𝗹𝗱 𝗮 𝘀𝗼𝗹𝗶𝗱 𝗱𝗮𝘁𝗮 𝗶𝗻𝗳𝗿𝗮𝘀𝘁𝗿𝘂𝗰𝘁𝘂𝗿𝗲 𝗮𝗻𝗱 𝗲𝗻𝗳𝗼𝗿𝗰𝗲 𝘀𝘁𝗿𝗶𝗰𝘁 𝗱𝗮𝘁𝗮 𝗵𝘆𝗴𝗶𝗲𝗻𝗲. Getting your house in order is the foundation for delivering on any AI ambition. The MIT Technology Review — based on insights from 205 C-level executives and data leaders — lays it out clearly: 𝗠𝗼𝘀𝘁 𝗰𝗼𝗺𝗽𝗮𝗻𝗶𝗲𝘀 𝗱𝗼 𝗻𝗼𝘁 𝗳𝗮𝗰𝗲 𝗮𝗻 𝗔𝗜 𝗽𝗿𝗼𝗯𝗹𝗲𝗺. 𝗧𝗵𝗲𝘆 𝗳𝗮𝗰𝗲 𝗰𝗵𝗮𝗹𝗹𝗲𝗻𝗴𝗲𝘀 𝗶𝗻 𝗱𝗮𝘁𝗮 𝗾𝘂𝗮𝗹𝗶𝘁𝘆, 𝗶𝗻𝗳𝗿𝗮𝘀𝘁𝗿𝘂𝗰𝘁𝘂𝗿𝗲, 𝗮𝗻𝗱 𝗿𝗶𝘀𝗸 𝗺𝗮𝗻𝗮𝗴𝗲𝗺𝗲𝗻𝘁. Therefore, many firms are still stuck in pilots, not production. Changing that requires strong data foundations, scalable architectures, trusted partners, and a shift in how companies think about creating real value with AI. Because pilots are easy, BUT scaling AI across the enterprise is hard. 𝗛𝗲𝗿𝗲 𝗮𝗿𝗲 𝘁𝗵𝗲 𝗸𝗲𝘆 𝘁𝗮𝗸𝗲𝗮𝘄𝗮𝘆𝘀: ⬇️ 1. 95% 𝗼𝗳 𝗰𝗼𝗺𝗽𝗮𝗻𝗶𝗲𝘀 𝗮𝗿𝗲 𝘂𝘀𝗶𝗻𝗴 𝗔𝗜 — 𝗯𝘂𝘁 76% 𝗮𝗿𝗲 𝘀𝘁𝘂𝗰𝗸 𝗮𝘁 𝗷𝘂𝘀𝘁 1–3 𝘂𝘀𝗲 𝗰𝗮𝘀𝗲𝘀: ➜ The gap between ambition and execution is huge. Scaling AI across the full business will define competitive advantage over the next 24 months. 2. 𝗗𝗮𝘁𝗮 𝗾𝘂𝗮𝗹𝗶𝘁𝘆 𝗮𝗻𝗱 𝗹𝗶𝗾𝘂𝗶𝗱𝗶𝘁𝘆 𝗮𝗿𝗲 𝘁𝗵𝗲 𝗿𝗲𝗮𝗹 𝗯𝗼𝘁𝘁𝗹𝗲𝗻𝗲𝗰𝗸𝘀: ➜ Without curated, accessible, and trusted data, no AI strategy can succeed — no matter how powerful the models are. 3. 𝗚𝗼𝘃𝗲𝗿𝗻𝗮𝗻𝗰𝗲, 𝘀𝗲𝗰𝘂𝗿𝗶𝘁𝘆, 𝗮𝗻𝗱 𝗽𝗿𝗶𝘃𝗮𝗰𝘆 𝗮𝗿𝗲 𝘀𝗹𝗼𝘄𝗶𝗻𝗴 𝗔𝗜 𝗱𝗲𝗽𝗹𝗼𝘆𝗺𝗲𝗻𝘁 — 𝗮𝗻𝗱 𝘁𝗵𝗮𝘁 𝗶𝘀 𝗮 𝗴𝗼𝗼𝗱 𝘁𝗵𝗶𝗻𝗴: ➜ 98% of executives say they would rather be safe than first. Trust, not speed, will win in the next AI wave. 4. 𝗦𝗽𝗲𝗰𝗶𝗮𝗹𝗶𝘇𝗲𝗱, 𝗯𝘂𝘀𝗶𝗻𝗲𝘀𝘀-𝘀𝗽𝗲𝗰𝗶𝗳𝗶𝗰 𝗔𝗜 𝘂𝘀𝗲 𝗰𝗮𝘀𝗲𝘀 𝘄𝗶𝗹𝗹 𝗱𝗿𝗶𝘃𝗲 𝘁𝗵𝗲 𝗺𝗼𝘀𝘁 𝘃𝗮𝗹𝘂𝗲: ➜ Generic generative AI (chatbots, text generation) is table stakes. True differentiation will come from custom, domain-specific applications. 5. 𝗟𝗲𝗴𝗮𝗰𝘆 𝘀𝘆𝘀𝘁𝗲𝗺𝘀 𝗮𝗿𝗲 𝗮 𝗺𝗮𝗷𝗼𝗿 𝗱𝗿𝗮𝗴 𝗼𝗻 𝗔𝗜 𝗮𝗺𝗯𝗶𝘁𝗶𝗼𝗻𝘀: ➜ Firms sitting on fragmented, outdated infrastructure are finding that retrofitting AI into legacy systems is often more costly than building new foundations. 6. 𝗖𝗼𝘀𝘁 𝗿𝗲𝗮𝗹𝗶𝘁𝗶𝗲𝘀 𝗮𝗿𝗲 𝗵𝗶𝘁𝘁𝗶𝗻𝗴 𝗵𝗮𝗿𝗱: ➜ From GPUs to energy bills, AI is not cheap — and mid-sized companies face the biggest barriers. Smart firms are building realistic ROI models that go beyond hype. 𝗕𝘂𝗶𝗹𝗱𝗶𝗻𝗴 𝗮 𝗳𝘂𝘁𝘂𝗿𝗲-𝗿𝗲𝗮𝗱𝘆 𝗔��� 𝗲𝗻𝘁𝗲𝗿𝗽𝗿𝗶𝘀𝗲 𝗶𝘀𝗻’𝘁 𝗮𝗯𝗼𝘂𝘁 𝗰𝗵𝗮𝘀𝗶𝗻𝗴 𝘁𝗵𝗲 𝗻𝗲𝘅𝘁 𝗺𝗼𝗱𝗲𝗹 𝗿𝗲𝗹𝗲𝗮𝘀𝗲. 𝗜𝘁’𝘀 𝗮𝗯𝗼𝘂𝘁 𝘀𝗼𝗹𝘃𝗶𝗻𝗴 𝘁𝗵𝗲 𝗵𝗮𝗿𝗱 𝗽𝗿𝗼𝗯𝗹𝗲𝗺𝘀 — 𝗱𝗮𝘁𝗮, 𝗶𝗻𝗳𝗿𝗮𝘀𝘁𝗿𝘂𝗰𝘁𝘂𝗿𝗲, 𝗴𝗼𝘃𝗲𝗿𝗻𝗮𝗻𝗰𝗲, 𝗮𝗻𝗱 𝗥𝗢𝗜 — 𝘁𝗼𝗱𝗮𝘆.
Data Quality for AI
Explore top LinkedIn content from expert professionals.
-
-
The AI Wave Finally Hit - and AI Data Readiness will be next. Over the past few weeks, something has shifted. OpenClaw has been making waves in the open-source world. Claude Cowork's plugins landed with a thud, particularly the automated contract-analysis tool. In days, hundreds of billions in market value were wiped off established software and IT services stocks. Markets don't reprice like that over a single product. They reprice when a deeper assumption breaks. 🔵 The Threshold Those of us working closely with systems like Claude Code have seen this coming - especially since Opus-class models made agentic workflows viable. But this feels like the moment the conversation crossed a threshold. What was niche is now mainstream. Here is a prediction: people will soon wake up to the importance of how these systems are grounded. Frameworks like OpenClaw are powerful, but rely on emergent behaviour over loosely structured context. An ontology-backed data structure gives you something tighter: clearer constraints, more predictable reasoning, and far less ambiguity about what the system is allowed to conclude. That difference shows up as reliability, and it will become impossible to ignore once people start to engage seriously. 🔵 A Personal Resonance For years, my argument was simple: AI is coming, and organisations need to get their data ready. It was never about chasing the latest technology. It was about recognising that once AI arrived, the limiting factor would not be the models - it would be the data. What I didn't anticipate is how unprepared I would be when that moment truly arrived. Last week, as I fully "wire-headed" into one of our internal agents - with direct access to our ontology and knowledge graph - something crystallised. The speed with which organisational context became usable, the way complex structures turned into leverage, was both exhilarating and unsettling. We are not psychologically prepared for what this is going to feel like. You can see this by the slightly manic look in the eyes of those who have already wire-headed. 🔵 The Principles That Still Hold As things get increasingly volatile, it's worth returning to core principles I've been repeating for a decade. First, focus on your data. AI is like an iceberg: what you see above the surface gets the attention, but what matters is what sits underneath. Second, "getting your data ready" means two things: linking it together richly, and organising it semantically. Without that, AI systems either underperform or produce confident nonsense. Finally, stick to open standards. They are the only reliable way to maintain flexibility as tools, vendors, and architectures change faster than organisations can react. The recent market reaction wasn't panic over a single tool. It was a delayed recognition of a reality building for years. AI didn't arrive overnight. But now that it's here, the cost of not being data-ready will become visible - all at once.
-
I invited 31 researchers to test AI research synthesis by running the exact same prompt. They learned LLM analysis is overhyped, but evaluating it is something you can do yourself. Last month I ran an #AI for #userresearch workshop with Rosenfeld Media. Our first cohort was full of smart, thoughtful researchers (if you participated in the workshop, I hope you’ll tag yourself and weigh in in the comments!). A major limitation of a lot of AI for UXR “thought leadership” right now is that too much of it is anecdotal: researchers run datasets a few times through a commercial tool and decide whether or not the output is good enough based on only a handful of results. But for nondeterministic systems like generative AI, repeated testing under controlled conditions is the only way to know how well they actually work. So that’s what we did in the workshop. Our workshop participants produced a lot of interesting findings about qualitative research synthesis with AI: 1️⃣ LLMs can product vastly different output even with the exact same prompt and data. The number of themes alone ranged from 5 to 18, with a median of about 10.5. 2️⃣ Our AI-generated themes mapped pretty well to human-generated themes, but there were some notable differences. This led to a discussion of whether mapping to human themes is even the right metric to use to evaluate AI synthesis (how are we evaluating whether the human-generated themes were right in the first place?). 3️⃣ The bigger concern for the researchers in the workshop was the lack of supporting evidence for themes. The supporting quotes the LLM provided looked okay superficially, but on closer investigation *every single participant* found examples of data being misquoted or entirely fabricated. One person commented that validating the output was ultimately more work than performing the analysis themselves. Now, I want to acknowledge that this is one dataset, one prompt (although, a carefully vetted one, written by an industry expert), and one model (GPT 4o 2024-11-20). Some researchers claim that GPT 4o is worse for research hallucinations–and perhaps it is–but it is still a heavily utilized model in current off-the-shelf AI research tools (and if you’re using off-the-shelf tools, you won’t always know which models they’re using unless you read a whole lot of fine print). But the point is–I think this is exactly the level at which we should be scrutinizing the output of *all* LLMs in research. AI absolutely has its place in the modern researcher’s toolkit. But until we systematically evaluate its strengths and weaknesses, we're rolling the dice every time we use it. We'll be running a second round of my workshop in June as part of Rosenfeld Media’s Designing with AI conference (ticket prices go up tomorrow; register with code PAINE-DWAI2025 for a discount). Or, to hear about other upcoming workshops and events from me, sign up for my mailing list (links below).
-
Every customer conversation I have about AI ends up coming back to data. Why? Because AI has raised the stakes on data quality. Messy data has always been a problem. But now, it’s a bigger one. When agents are making decisions and taking actions on top of bad data, you don't just get bad outputs. You get bad emails triggered to customers. Bad deal intelligence. And errors that compound across workflows moving at AI speed. Agents need data that is accurate, up to date, and unified. But there's no magic wand for getting there. AI can make querying easier. It can make joins better. But if your data lives in silos, definitions aren't clear, and quality is poor, AI doesn't fix that. It inherits it. Take a simple term like ARR (Annual Recurring Revenue). Finance defines it one way. Sales defines it another. Both are legitimate. They're just not the same number. Before any agent touches that data, a human who understands the business has to decide which definition should apply. And if you miss that step, AI will confidently hand you the wrong answer every time. Teams need to invest the time in getting to clean data, clear definitions, and consistent labels. Doing this work upfront is what enables outcomes down the line. And it's where so many companies are falling short. So for any business leader looking to be successful with AI, focus on your data readiness first. We have an entire ecosystem of partners who can help you get started.
-
For years, transformation focused on digitization. Today, we see a move toward continuous reinvention, where Ai connects strategy, operations, and talent in new ways. Yet without a strong data foundation, that vision remains out of reach. Ai-ready data changes the equation. It brings consistency, context, and trust into every decision. It allows organizations to move from isolated insights to enterprise-wide intelligence. At scale, this creates a different kind of company, one that learns faster and acts with precision. The organizations that will lead are those that take a disciplined approach. They define clear ownership of data. They embed governance into everyday processes. They modernize architecture to support both scale and flexibility. And they build a culture where data informs every level of decision-making. Ai will continue to advance. That is certain. The real differentiator will be how well organizations prepare their data to keep pace. Reinvention starts there. https://lnkd.in/gktFvYYj Accenture
-
You wouldn't cook a meal with rotten ingredients, right? Yet, businesses pump messy data into AI models daily— ..and wonder why their insights taste off. Without quality, even the most advanced systems churn unreliable insights. Let’s talk simple — how do we make sure our “ingredients” stay fresh? Start Smart → Know what matters: Identify your critical data (customer IDs, revenue, transactions) → Pick your battles: Monitor high-impact tables first, not everything at once Build the Guardrails: → Set clear rules: Is data arriving on time? Is anything missing? Are formats consistent? → Automate checks: Embed validations in your pipelines (Airflow, Prefect) to catch issues before they spread → Test in slices: Check daily or weekly chunks first—spot problems early, fix them fast Stay Alert (But Not Overwhelmed): → Tune your alarms: Too many false alerts = team burnout. Adjust thresholds to match real patterns → Build dashboards: Visual KPIs help everyone see what's healthy and what's breaking Fix It Right: → Dig into logs when things break—schema changes? Missing files? → Refresh everything downstream: Fix the source, then update dependent dashboards and reports → Validate your fix: Rerun checks, confirm KPIs improve before moving on Now, in the era of AI, data quality deserves even sharper focus. Models amplify what data feeds them — they can’t fix your bad ingredients. → Garbage in = hallucinations out. LLMs amplify bad data exponentially → Bias detection starts with clean, representative datasets → Automate quality checks using AI itself—anomaly detection, schema drift monitoring → Version your data like code: Track lineage, changes, and rollback when needed Here's the amazing step-by-step guide curated by DQOps - Piotr Czarnas to deep dive in the fundamentals of Data Quality. Clean data isn’t a process — it’s a discipline. 💬 What's your biggest data quality challenge right now?
-
13 national cyber agencies from around the world, led by #ACSC, have collaborated on a guide for secure use of a range of "AI" technologies, and it is definitely worth a read! "Engaging with Artificial Intelligence" was written with collaboration from Australian Cyber Security Centre, along with the Cybersecurity and Infrastructure Security Agency (#CISA), FBI, NSA, NCSC-UK, CCCS, NCSC-NZ, CERT NZ, BSI, INCD, NISC, NCSC-NO, CSA, and SNCC, so you would expect this to be a tome, but it's only 15 pages! It is refreshing to see that the article is not solely focused on LLMs (eg. ChatGPT), but defines Artificial Intelligence to include Machine Learning, Natural Language Processing, and Generative AI (LLMs), while acknowledging there are other sub-fields as well. The challenges identified (with actual real-world examples!) are: 🚩 Data Poisoning of an AI Model: manipulating an AI model's training data, leading to incorrect, biased, or malicious outputs 🚩 Input Manipulation Attacks: includes prompt injection and adversarial examples, where malicious inputs are used to hijack AI model outputs or cause misclassifications 🚩 Generative AI Hallucinations: generating inaccurate or factually incorrect information 🚩 Privacy and Intellectual Property Concerns: challenges in ensuring the security of sensitive data, including personal and intellectual property, within AI systems 🚩 Model Stealing Attack: creating replicas of AI models using the outputs of existing systems, raising intellectual property and privacy issues The suggested mitigations include generic (but useful!) cybersecurity advice as well as AI-specific advice: 🔐 Implement cyber security frameworks 🔐 Assess privacy and data protection impact 🔐 Enforce phishing-resistant multi-factor authentication 🔐 Manage privileged access on a need-to-know basis 🔐 Maintain backups of AI models and training data 🔐 Conduct trials for AI systems 🔐 Use secure-by-design principles and evaluate supply chains 🔐 Understand AI system limitations 🔐 Ensure qualified staff manage AI systems 🔐 Perform regular health checks and manage data drift 🔐 Implement logging and monitoring for AI systems 🔐 Develop an incident response plan for AI systems This guide is a great practical resource for users of AI systems. I would interested to know if there are any incident response plans specifically written for AI systems - are there any available from a reputable source?
-
#UltrasoundAI didn’t struggle because the models were weak. It struggled because the images were. For years, we kept asking smarter algorithms to interpret noisy, operator-dependent ultrasound data and then blamed AI when results didn’t translate clinically. I was proud to be part of a newly published, peer-reviewed study with the Mayo Clinic that took a different approach: Fix the data foundation first. Using over 62,000 real-world breast ultrasound scans, we evaluated what happens when raw B-mode images are combined with quality-improved and enhanced ultrasound representations before AI ever sees them. Here’s what stood out: • Same patient. Same scan. Different outcome. Multifeature ultrasound improved diagnostic accuracy by +66% versus B-mode alone. • One size doesn’t fit all in medicine and that’s a good thing. • Graph models prioritized sensitivity (91.7%) for screening • Masked autoencoders prioritized specificity (100%) for diagnosis • The breakthrough wasn’t “better AI.” It was giving AI clearer, more consistent signal to work with. • This matters beyond academic centers. Image quality has been the limiting factor for AI at the point of care. Improve the signal, and AI becomes usable where patients actually are. This work reinforced something clinicians already know intuitively: You can’t out-optimize poor inputs. If we want AI diagnostics to scale safely, equitably, and responsibly, image quality has to be treated as infrastructure not an afterthought. Full disclosure: I serve as an advisor to #PONSAI, and it’s been encouraging to work alongside leaders like soner ozkan and Ilker Hacihaliloglu who are focused on strengthening the foundations of clinical AI, not just chasing benchmarks. Where do you see AI breaking down most today the algorithm, or the data we’re feeding it? #HealthcareAI #MedicalImaging #Ultrasound #ClinicalAI #DigitalHealth #AIinMedicine #PointOfCare #HealthTech #AIValidation Eric Topol, MD this ironically your post today.
-
The next AI bottleneck is not model intelligence. It is data readiness. We already know models can reason, write, search, summarize, and call tools. The harder problem now is what happens when those models are connected to enterprise data. Because in production, AI does not fail only because the model is weak. It fails when the data is: stale duplicated incomplete poorly governed spread across too many systems missing context not trusted by the business And once AI starts taking action, bad data does not just create bad dashboards. It creates bad decisions at machine speed. That is why the conversation around enterprise AI is shifting from “Which model should we use?” to: “Is our data ready for AI to act on it?” Informatica World 2026 on Salesforce+ has some useful on-demand sessions around AI-ready data, trusted data foundations, governance, and the future of enterprise data management. For architects, data leaders, and AI builders, this is the layer that matters most. Because agentic AI will not scale on top of messy data. It will scale on top of trusted data. Watch it here: https://lnkd.in/dCDDMFuX
-
📘 Downscaling CHIRPS Precipitation Data to 100m Resolution Using Sentinel-2 in Google Earth Engine Source Code = https://lnkd.in/dNc7NbjE 1. Introduction: Rainfall data at high spatial resolution is critical for precise hydrological analysis, drought monitoring, and agriculture planning. However, most global precipitation datasets, such as CHIRPS (Climate Hazards Group InfraRed Precipitation with Station data), are available at coarser resolutions (~5 km). This project addresses this limitation by downscaling CHIRPS daily precipitation data to 100-meter spatial resolution using bilinear interpolation and Sentinel-2 as a high-resolution spatial reference. 2. Objective: To extract and sum CHIRPS precipitation data over a selected AOI (WMH District) for a specific 3-month period (October 2023 – January 2024). To downscale the CHIRPS raster data to a finer 100-meter resolution using Sentinel-2 spatial referencing. To visualize and compare the original and downscaled precipitation maps. To prepare refined precipitation layers for potential integration with NDVI, crop condition analysis, or drought indices. 3. Importance of the Study: Higher spatial resolution enables more localized analysis of rainfall, especially in heterogeneous landscapes. Improved input for climate models and agro-hydrological studies. Better decision-making for irrigation scheduling, water resource management, and drought preparedness. Supports integration with high-resolution datasets such as NDVI, land use, or soil moisture for multi-parameter environmental studies. 4. Benefits: Enhanced accuracy in rainfall data analysis at local and district levels. Scalable method applicable to any region globally. Supports policymakers and researchers with higher-resolution inputs for climate resilience, agricultural planning, and hydrological monitoring. Efficient use of cloud computing via Google Earth Engine for handling large spatiotemporal datasets. 5. Output: Original CHIRPS Precipitation Map (Oct 2023 – Jan 2024) clipped to WMH District. Downscaled Precipitation Map at 100m resolution, reprojected using Sentinel-2 reference. Color-coded visualization using a 5-class blue gradient, where: Light blue = Low precipitation Dark blue = High precipitation Ready-to-export raster layer of downscaled precipitation (if export added). Output maps can be further used for vegetation correlation (e.g., NDVI vs. rainfall) and SPI generation. #GEE #GoogleEarthEngine #BuildupAreaExpansion #GeospatialAnalytics #RemoteSensing #UrbanExpansion #Geospatial #GoogleEarthEngine #GIS #SustainableDevelopment #Sentinel2 #GeospatialTech #PhD #Agriculture #ClimateSmart #GIS #DeepLearning #ClimateSmartAgriculture #CropHealthMonitoring #DroughtMonitoring #SustainableFarming #Sentinel2 #GoogleEarthEngine #NDVI #LandsatData #GISMapping #GeospatialAnalysis #AIinAgriculture #EarthObservation #AgricultureMapping #RemoteSensin #SatelliteImagery