Insights

What data does a non-technical organisation need before it starts with AI?

“Our data is not ready” is the sentence that stops more AI conversations in Mauritian boardrooms than any other. Sometimes it is true. Often it is a polite way of saying that nobody knows where the data is, who is responsible for it, or whether the organisation is allowed to use it for something new.

This article is written for a managing director or board member who has no data team and no intention of building one. It explains what “the data can feed a model” actually means for the four AI uses most organisations consider first, gives a data inventory method that takes days rather than months, and helps you tell the difference between data that is a genuine blocker and data that has become an excuse.

Why the data question is asked badly

The question “is our data ready for AI?” has no single answer because “AI” is not one thing. A generative AI tool that drafts a letter needs almost nothing from your systems. A forecasting model needs years of clean, consistent history. Treating both as the same problem leads to two failures: organisations delay low-risk uses that need no special data, and they rush into data-hungry uses that were never going to work.

RAND’s 2024 study of why AI projects fail, based on interviews with 65 experienced data scientists and engineers, found that the leading root cause is that stakeholders misunderstand or miscommunicate what problem needs to be solved, and that the second is that the organisation lacks the data needed to train an effective model. One interviewee put it plainly: leaders think they have great data because they receive weekly sales reports, without realising that the data behind those reports may not serve a new purpose.

Both findings point in the same direction. Data readiness can only be assessed against a specific, well-defined use.

What “the data can feed a model” means for four common uses

1. Document and email drafting with generative AI

This is the lowest data bar of the four. Tools such as ChatGPT, Copilot or Gemini draft from the instructions you give them and the material you paste in. They do not need access to your databases. What they need is examples of what “good” looks like in your organisation: a few well-written proposals, your standard letter formats, your tone.

The data question here is not “do we have enough?” but “what are we allowed to put in?” A staff member pasting a customer’s complaint, with the customer’s name and account details, into a public tool is the real risk. That is a policy question, not a volume question.

Minimum condition: a small library of approved reference documents and a clear rule on what may not be entered into a public tool.

2. Customer service assistants

An assistant that answers customer questions needs a knowledge base to answer from: product descriptions, price lists, terms and conditions, FAQs, policies, procedures. It is only as accurate as those documents are current and consistent.

The typical problem is not absence but contradiction. The website says one thing, the printed brochure another, and the customer service team has a third version in their heads. An assistant trained on all three will confidently give the wrong answer.

If the assistant is expected to look up a specific customer’s order or account, the bar rises sharply. It now needs a live connection to your systems, personal data is involved, and the Data Protection Act 2017 applies in full.

Minimum condition: one owned, current, non-contradictory set of reference documents, and a decision on whether the assistant will touch personal data at all.

3. Forecasting and analytics

This is where “data readiness” earns its reputation. A model that predicts demand, cash flow, stock levels or churn learns from history. It needs enough past examples, recorded consistently, with the fields that matter filled in the same way over time.

Monthly sales forecasting generally needs at least two or three full years so that seasonality appears more than once. Consistency matters as much as volume: if the product codes changed in 2024, or a field was left blank for six months while a system migrated, the history is broken at that point. Ownership matters too. The finance system, the point-of-sale system and the sales manager’s spreadsheet are usually three different versions of the truth, and a forecasting project starts by deciding which one is authoritative.

Minimum condition: two or more years of consistent history in one authoritative source, with a named owner who can explain what each field means and when definitions changed.

4. Process automation

Automation, whether workflow tools or newer AI agents, needs structured inputs. If invoices arrive as PDFs in an inbox, an AI tool can read them, but only if the invoices are reasonably consistent and there is a clear place for the extracted data to go. If approvals happen on WhatsApp and the record is a screenshot, the process is not automatable until the process itself is fixed.

Minimum condition: documented process steps, consistent input formats, and a system where the output can be recorded.

How different these four really are

Use Needs history? Needs a live system connection? Touches personal data? Main data risk
Drafting with generative AI No No Only if staff paste it in Confidential or personal data entered into public tools
Customer service assistant No (needs current documents) Only if it looks up accounts Often Contradictory or outdated reference documents
Forecasting and analytics Yes, two or more years Usually Sometimes Inconsistent or broken history, unclear ownership
Process automation No Yes Often Inconsistent inputs, undocumented rules

The practical lesson: a non-technical organisation can usually start with the first use and often the second, while the third and fourth are where the data inventory below becomes essential.

The five minimum conditions, in plain language

Whatever the use, five questions determine whether a dataset can be relied on. They are the questions Consultaix asks in a readiness assessment, and none of them requires a technical answer.

Where does it live? Name the system, file or person. “In the accounting software” is an answer. “Somewhere in shared drives” is not. If the same data lives in two places, note both and decide which is authoritative.

Who owns it? Not who typed it in, but who is accountable for it being correct and who can authorise its use. In most SMEs this is a head of department. If nobody can be named, the data is not owned, and unowned data is not ready.

Is it consistent? Are the same things recorded the same way over time and across branches? Ask the person who uses it daily: “if I compared 2023 to 2025, would the categories match?” They will know.

Are we allowed to use it? Personal data collected for one purpose cannot simply be repurposed. The Data Protection Act 2017 requires that personal data are collected for explicit, specified and legitimate purposes and not further processed in a manner incompatible with those purposes (section 21). Using customer records to train a model is a new purpose, and someone needs to check whether it fits.

Is there enough history? Only relevant for forecasting and analytics, but decisive there. Count the months of consistent, complete records. If the answer is fewer than 24, forecasting is probably premature.

A data inventory that takes days, not months

A full data audit is a project. A data inventory sufficient to decide whether to proceed with a specific AI use is a few days of a capable manager’s time. The method is deliberately simple.

Day 1. Define the use. Write one sentence: “We want AI to do X, using Y, so that Z.” If you cannot write the sentence, stop. This is the problem definition failure RAND identifies, and no amount of data work fixes it.

Day 1 to 2. List the data the use would need. Work backwards from the sentence. A customer service assistant needs product information, policies and FAQs. A cash flow forecast needs receipts, payments, invoices and their dates. Keep the list to the ten or fifteen items that matter.

Day 2 to 3. Fill in the table. For each item, answer the five questions above. The department head who works with the data can usually answer in a twenty-minute conversation. Do not ask IT to run reports at this stage; ask people what they know.

Day 4. Decide. Mark each row green, amber or red. Any red row on “allowed to use” is a hard stop until resolved. Amber rows on ownership or consistency are workable if there is a named person willing to fix them. Then decide: proceed, proceed with a narrower scope, or fix first.

The inventory template

Data item Where it lives Owner (name and role) Consistent over time? Allowed for this use? History available Status
Customer master list CRM (system name) Head of Sales Yes since 2023; before that, different format Personal data; purpose to be confirmed 5 years Amber
Product catalogue and prices Website and Excel (differ) Nobody named No, two versions Yes, no personal data Not applicable Amber
Monthly sales by product Accounting system Finance Manager Product codes changed Jan 2024 Yes, aggregated 3 years, break in 2024 Amber
Complaint emails Shared inbox Customer Service Lead Free text, no categories Personal data; consent unclear 2 years Red
Standard operating procedures Shared drive, some printed Operations Manager Several out of date Yes, no personal data Not applicable Amber

The rows above are illustrative. The pattern they show is typical: most of the data exists, most of it has an owner who can be found, and the real work is in the “consistent” and “allowed” columns.

Five signs data is the real blocker

Data is genuinely the blocker, and the AI use should wait, when:

  1. The history has a break that cannot be reconciled. A system migration, a change of product coding or a reorganisation means the “before” and “after” cannot be compared, and the use needs the full period.
  2. Nobody can be named as owner. If the finance manager says it is the sales team’s data and the sales manager says it is finance’s, the data will not be maintained and the model will decay.
  3. The authoritative source does not exist. Three spreadsheets disagree and nobody can say which is right. A model built on one of them will be challenged the moment it produces an uncomfortable result.
  4. The legal basis for the new use is unclear and the use depends on personal data. If the use cannot proceed without customer or staff records, and nobody can say under which of the lawful bases in section 28 of the Data Protection Act 2017 the processing would fall, this must be resolved first.
  5. The volume is genuinely too small. Twelve months of records for a seasonal business, or forty customers for a churn model, is not enough for statistical methods, however clean.

Five signs data is an excuse

Data is being used as a reason to delay, when the actual reasons lie elsewhere, if:

  1. The proposed use does not need historical data. Drafting, summarising and document assistants need reference material and rules, not databases. “Our data is not ready” is irrelevant to them.
  2. The concern is vague. “Data quality” without a named field, system or example is a worry, not a finding. Ask which specific data item is unusable and for which use.
  3. Nobody has actually looked. If the inventory above has not been attempted, the readiness of the data is unknown, not poor. Unknown is fixable in days.
  4. The fix is a decision, not a project. Two price lists that disagree are resolved by a manager choosing one. Undefined ownership is resolved by naming someone. These take a meeting, not a budget.
  5. The same objection stops every initiative. If data readiness has blocked the CRM upgrade, the dashboard project and now AI, the organisation has a decision-making problem that it is describing as a data problem.

The honest test is this: could a specific, named person fix the specific, named issue within a month? If yes, it is a task. If no, it is a blocker.

Where data sits on the Consultaix AI Adoption Ladder

Consultaix assesses readiness across five dimensions: strategy, governance, data, capability and execution. Each is scored on the Consultaix AI Adoption Ladder, which has four rungs. For the data dimension, the rungs are described as follows.

Rung What the data dimension looks like
Aware Data has not been inventoried. The organisation knows it holds information but cannot say where, who owns it, or whether it can be used.
Exploring Data has been located. Ownership, quality and consent for new uses are still unknown or partially known.
Adopting The live AI use case has an owned, clean dataset. Someone is accountable for it, it is consistent enough for the purpose, and its use has been checked against the law.
Scaling Ownership, quality checks and access rules are standard across functions rather than confined to one use case. New AI uses can be assessed against existing rules rather than starting from scratch.

Two points deserve emphasis. First, the step from Aware to Exploring is the inventory described above, and it takes days. Second, Adopting is defined per use case. An organisation can be at Adopting for a customer service assistant and at Aware for forecasting, because the datasets are different. The other dimensions interact with data: governance determines who may decide on data use, and capability determines whether anyone can maintain the dataset once the project is live. A clean dataset with no owner will not stay clean.

The Data Protection Act 2017 as the boundary

For any use that involves personal data, meaning information about identifiable customers, staff, suppliers or members of the public, the Data Protection Act 2017 sets the boundary on what may be used and how. This is not optional and it is not a matter of good practice.

Legal requirement. Section 21 of the Act sets out the principles every controller or processor must follow: personal data must be processed lawfully, fairly and transparently; collected for explicit, specified and legitimate purposes and not further processed in a manner incompatible with those purposes; adequate, relevant and limited to what is necessary; accurate and, where necessary, kept up to date. For AI readiness, the purpose principle is the one most often overlooked. Data collected to fulfil an order was collected for that purpose. Using it to train a model is a different purpose, and the organisation must be able to justify the compatibility.

Legal requirement. Section 28 lists the lawful bases for processing: the data subject’s consent for specified purposes, or processing necessary for a contract, a legal obligation, vital interests, public tasks, the controller’s legitimate interests (subject to a harm test), or historical, statistical or scientific research. Before any personal data feeds an AI system, the organisation should be able to point to which basis applies.

Legal requirement. Section 22 requires every controller to adopt policies and implement appropriate technical and organisational measures, including keeping a record of processing operations and performing a data protection impact assessment where section 34 requires it. Section 31 requires appropriate security measures against unauthorised access, alteration, disclosure or loss.

Best practice. Run the “allowed to use” column of the inventory with the Act’s principles in front of you. Where a dataset contains personal data and the AI use is new, treat it as amber at best until someone has recorded the lawful basis and considered whether an impact assessment is needed. The Data Protection Office publishes guidance on both, linked in the sources below.

Consultaix recommendation. For a first AI use, prefer datasets that contain no personal data at all: product information, procedures, policies, aggregated figures. This removes the hardest column from the inventory and lets the organisation build experience before it takes on the compliance work that personal data requires.

What to do this month

Write the one-sentence use. List the data items it needs. Spend three days filling in the inventory table with the department heads and mark the rows. If any row is red on “allowed to use”, resolve that first. If the rest are green or amber with a named owner, the data is ready enough for a pilot, and the remaining issues become part of the project rather than a reason not to start it.

If the exercise reveals that nobody can name an owner for anything, the finding is still valuable. The organisation has learned, in days, that it is at the Aware rung on data and that the first investment should be in ownership and consistency, not in tools.

Where Consultaix fits

An AI Readiness Assessment scores the data dimension alongside strategy, governance, capability and execution on the Consultaix AI Adoption Ladder, so the data inventory is done once and in context rather than repeated for every vendor conversation. Where the assessment shows the data is ready enough for a defined use, AI Implementation and Automation takes the inventory as its starting point. The AI Adoption Ladder page describes all four rungs across the five dimensions.

Sources