Fundamentals GUIDE

AI & Data

Data is the recorded information a machine-learning system learns from or processes.

  • 2 min read
  • Last updated
On this page2 min read
  1. Overview
  2. Key takeaways
  3. Deep Dive
  4. Find leakage in a cancellation dataset
  5. Strategic Impact
  6. Real-World Implementation
  7. Risks & Guardrails
  8. Implementation Roadmap
  9. Sources and further reading
  10. Keep Exploring
  11. Frequently asked questions

Overview

Its usefulness depends on relevance, measurement quality, permissions, and coverage of the intended task. More records do not automatically correct systematic errors or missing populations.

Key takeaways

  1. Define the unit of an example.
  2. Use only information available at prediction time.
  3. Track data provenance, missingness, and subgroup coverage.

Deep Dive

Start by defining what one example represents. A row might describe a customer, a transaction, a photograph, or one moment in a time series. Those units determine how duplicates, labels, and evaluation splits should work. Ten measurements from one device are not necessarily ten independent devices.

Features are inputs available to the model. Labels are target outcomes used in supervised learning. Check when each feature becomes available: a cancellation reason recorded after a customer leaves cannot fairly predict that departure beforehand. This is a form of leakage even when the field looks highly predictive.

Inspect missing values, annotation disagreements, unusual ranges, and changes in collection methods. Missing information can carry meaning; replacing every missing value with zero can conflate an unknown quantity with a real zero. Document the treatment and test it on representative examples.

Record provenance and access rules alongside the dataset. A public URL alone does not establish permission to reuse every item for every purpose. Collect only information needed for the task and define retention and deletion procedures. Evaluate separately on groups or conditions where errors would otherwise disappear inside an overall average.

04Worked example

Find leakage in a cancellation dataset

  1. Imagine records with signup date, monthly usage, cancellation date, and cancellation reason.

  2. To predict cancellations at the start of June, freeze every input at that date. Remove reasons and dates recorded after the prediction time.

  3. Train on earlier periods and test on a later untouched period. Compare results with and without the leaked fields.

What it shows

This hypothetical design exercise identifies an invalid shortcut before a flattering score becomes a deployment decision.

Strategic Impact

Clearer decisions

It helps you separate clear technical claims from marketing language.

Cost and budget

You can ask better implementation questions before spending money or time.

Team and workflow

Teams with shared understanding make better product, policy, and learning decisions.

Real-World Implementation

Separate multiple photographs of the same object before splitting a recognition dataset.

Flag a sensor reading outside the physically plausible range for review.

Risks & Guardrails

  • Different teams may use the same term differently, so define scope early.

  • Benchmarks can look strong while real-world performance is uneven.

  • Ignoring data quality and evaluation plans often creates fragile outcomes.

Implementation Roadmap

  1. Start with a plain-language definition of the outcome you need.

  2. Pick one success metric and one failure condition before testing.

  3. Run a small pilot with representative data, not a polished demo set.

  4. Document where AI & Data helps and where simpler methods are better.

Sources and further reading

  1. GoogleDataset characteristics

Keep Exploring

Free newsletter

Keep up with AI in 3 minutes a day

One short email each weekday with the three AI stories that actually matter. Free forever, no ads.

One email each weekday. Unsubscribe in one click. We never sell or share your address.

Test yourself

Take the AI & Data quiz

Instant feedback on every answer, and a shareable certificate with a verifiable ID once you pass a course.

Start quiz

Support free AI education. AI Understanding is a 501(c)(3) nonprofit — no ads, no paywall, ever. Make a donation

Frequently asked questions

Can a large dataset still be poor?

Yes. Duplicated, mislabeled, irrelevant, or systematically incomplete records can make a large dataset unsuitable for the intended task.