Most market research now begins on the open web. Pricing pages, product catalogs, marketplace listings, job postings, regulatory filings, and consumer reviews are public, current, and often more cost-efficient to observe at scale than commissioning equivalent surveys or panels. The difficulty is no longer finding the data but collecting it in a way that is both accurate and defensible. A dataset that is incomplete, geographically skewed, or gathered without regard for the rules is worse than no dataset at all, because it produces confident conclusions from a flawed foundation.
This is the part of the process that rarely makes it into the final report. Analysts describe their models, their sources, and their assumptions in detail, but the collection layer - how the raw observations were actually obtained - is usually compressed into a single line. That gap is where most data-quality and compliance risk lives. The following is a practical framework for closing it.
Start by Separating Public from Personal
The first distinction that governs everything downstream is not technical - it is legal. Publicly accessible data and personal data are two different categories, and they are treated very differently by regulators.
A product price, a stock level, a published specification, or a company's posted job listing is public, non-personal information. An individual's review, profile, or contact detail is personal data, and under frameworks such as the European Union's General Data Protection Regulation, personal data carries obligations around lawful basis, purpose limitation, and retention regardless of whether it happens to be visible on a public page. The European Commission's definition of personal data is deliberately broad, and "it was already public" is not, on its own, a lawful basis for processing it.
For most market-sizing, pricing, and competitive-intelligence work, the data of interest is non-personal - which keeps the project on far simpler ground. The discipline is to design collection so that personal data is excluded by default and captured only where there is a clear, documented reason and basis to do so.
Respect the Site's Stated Boundaries
The second layer is the publisher's own signals. The Robots Exclusion Protocol - the mechanism behind the familiar robots.txt file, formally standardized by the IETF in 2022 - lets a site declare which paths automated agents should not access. Honoring it is widely considered a signal of good faith, although it is not legally binding and practices vary across organizations and jurisdictions.
It is worth being precise about what robots.txt is and is not. It governs automated crawling conventions; it is not the full body of law on data access. In the U.S., the Ninth Circuit's rulings in hiQ Labs v. LinkedIn suggest that scraping publicly accessible data may not violate the Computer Fraud and Abuse Act under certain conditions. That interpretation is jurisdiction-specific, however, and does not eliminate risks related to contract law, data protection frameworks, or future litigation. In practice, contractual restrictions in a site's terms of service are one of the most common sources of legal exposure, even when the data itself is publicly accessible.
A commonly adopted, risk-minimizing posture is straightforward: collect only what is genuinely public, stay out of paths the site asks automated agents to avoid, never circumvent a login, and keep request volume low enough that it places no meaningful load on the source. The specific legal implications still vary by jurisdiction and context, and teams should validate their approach with legal counsel where appropriate.
Design for Representativeness, Not Just Volume
Once the boundaries are set, the question becomes one familiar to every researcher: is the sample representative?
Web data introduces a subtle sampling bias that surveys do not. What a website shows depends on who appears to be asking. Pricing, product availability, promotions, and even which currency or language is displayed are routinely tailored to the visitor's location. A pricing study run entirely from a single office IP in one country is not a global pricing study - it is a one-city snapshot that has been mislabeled. For example, pricing studies conducted from a single geographic location often misinterpret regional promotions or currency differences as global trends, producing systematically biased conclusions.
