How Data Is Generated and Collected#
š¦ Data Preparation 𧬠Data Types & Structure Lesson 001
Next ā¶ Ā· ā Section Ā· ā Hub
Important
⨠AI-generated content. This page was written with the assistance of an AI language model and is provided as a learning aid. Despite careful review, it may still contain mistakes, omissions, or out-of-date information. Whether you are new to the topic, a team lead, or a senior practitioner, treat it as a starting point rather than an authoritative reference: read it critically and independently verify anything you act on (code, commands, figures, and factual claims) against official documentation and primary sources before relying on it.
Where data comes from#
Section 2 treated data as something you have; this section starts one step earlier, with where it comes from. Understanding how data is generated and collected is the foundation of the Prepare phase, because the origin of a dataset determines what it can honestly be used for ā a point the context and bias lessons already foreshadowed.
How data is generated#
Data comes into existence in a few characteristic ways:
Observational ā recording what happens without intervening: transactions as they occur, clicks as users browse, sensor readings over time. Most business data is observational, and it shows what did happen, not necessarily what causes what.
Experimental ā deliberately varying something and measuring the result: the A/B test that shows two homepage designs to comparable groups. Experiments are what let analysis speak about causes rather than only associations.
Self-reported ā people telling you directly: surveys, forms, registrations. Rich and often the only route to why, but filtered through memory, honesty, and who chose to respond.
Derived ā computed from other data: a ācustomer lifetime valueā field built from transaction history. Convenient, but only as sound as its inputs and its formula.
Sources: first-, second-, and third-party#
Independently of how it is generated, data is classified by whose it is:
First-party ā collected by your own organisation directly from its own activity and customers. Usually the most trustworthy and relevant, because you control and understand its collection.
Second-party ā another organisationās first-party data, obtained directly from them through a partnership. Its quality depends on their collection practices, which you must ask about.
Third-party ā aggregated and sold by an entity that did not collect it from the original source. Broad and convenient, but its provenance and quality are the hardest to verify ā treat with corresponding caution.
The reliability gradient generally runs first ā second ā third-party, and it maps directly onto how much you can know about the collection context.
Why origin governs use#
Every downstream question about a dataset traces to its origin. Can this show causation? ā only if it was experimental. Does it represent all customers? ā only if collection reached them all. Can I trust the definitions? ā most where you controlled collection, least where a third party did. Knowing generation and source is how you answer these before, not after, building an analysis on the data.
The caveat#
Origin is often undocumented ā data arrives without a clear record of how or by whom it was collected, and reconstructing that is real detective work. When origin cannot be established, that uncertainty is itself a finding to state, not a detail to gloss: an analysis built on data of unknown provenance inherits unknown risk. The next lesson turns from where data comes from to which data a question actually needs.
Hint
See also
Source article Adapted (context, re-expressed) in our own words from: https://insightful-data-lab.com/2023/09/04/how-data-is-generated-and-collected/ (insightful-data-lab.com).