Skip to main content
← Back to blog

I Used Public Data to Find a Hidden Market (And You Can Too)

·13 min read

Introduction

Forty-two percent of startups fail because they build something nobody wants. That's the number from CB Insights that haunts me every time I start a new project. I used to think I could avoid that trap by talking to a few friends and reading some industry blogs. Spoiler: I was wrong. After wasting months on a product that flopped because I didn't really understand the market, I stumbled onto something that changed everything, using public data to uncover genuine demand. It's not as sexy as machine learning or AI, but it's a lot more reliable.

Here's the thing: most founders and sales teams make the same mistake. They assume they know who their customer is. They build a persona based on assumptions and then wonder why outreach falls flat. But the answer is already out there, in job boards, funding announcements, online communities, and yes, even in public records. You just need to know where to look and how to connect the dots. In this article, I'll walk you through the exact framework I now use to validate a market in weeks, not months.

The moment I realized my market research was broken

It was late 2023. I was consulting for a B2B SaaS startup that had built a collaboration tool for remote engineering teams. They'd raised a seed round based on a compelling story, but after six months of sales, they had exactly three paying customers, all of whom were friends of the CEO. The team was frustrated. They'd done "market research", surveys, competitor analysis, even a few dozen customer interviews. But somehow, they missed the real problem.

I decided to dig into the data. I pulled job postings from companies in their target verticals, scraped LinkedIn for roles like "remote engineering manager," and looked at funding announcements for startups in the collaboration space. What I found was sobering: the actual job postings rarely mentioned collaboration tools; the biggest pain point was onboarding and compliance. The product they built solved a problem that existed only in their imagination. The market didn't need it.

That experience taught me a lesson I'll never forget: market research based on assumptions is worse than no research at all, it gives you false confidence. The only way to uncover real demand is to look at what people actually do, not what they say they want. And the best source for that? Public data.

But wait, doesn't everyone already know this? Apparently not. A survey by Startup Genome found that 74% of startup failures are due to premature scaling, often rooted in faulty market assumptions. The same CB Insights report shows that understanding customer needs is the top factor for success. Yet most early-stage teams still rely on gut feelings and nice-to-have interviews. Why? Because digging through public data sounds tedious. But it doesn't have to be.

What most founders get wrong about demand

Let's be blunt: most people confuse "interesting problem" with "demand." Just because you can articulate a pain point doesn't mean people will pay to solve it. The classic mistake is to design a survey or interview question that leads the witness. "Would you use a tool that does X?" Of course they say yes, it's polite. But when it's time to open their wallet, the answer changes.

Another common error is relying on secondary market reports from big analyst firms. Those reports are useful for trends, but they often aggregate data that's too broad. A report about "AI in healthcare" doesn't tell you if radiologists in community hospitals will pay for your specific solution. You need granular signals, actual behavior.

And the third mistake? Ignoring the competition. Not just direct competitors, but the substitute behaviors that people use today. If you're building a new project management tool, your competition isn't just Asana, it's spreadsheets, sticky notes, and Slack messages. Public data can show you how many people are searching for those substitutes and what they complain about.

Consider this: a startup called Reflectly tried to build a wellness app for "happiness tracking." They surveyed users, and everyone said they'd love it. But when they launched, no one paid. Why? Because the substitute behavior (venting to a friend or journaling) was free and good enough. Public data from Reddit showed thousands of posts about "how to track my mood" but very few about "I wish there was an app." That gap between interest and action is where demand dies.

Step 1: Mining public data for market signals

The first step is to gather raw signals. You don't need expensive tools or subscriptions, just a browser and a spreadsheet. Here's where I start:

  • Job postings: Every job description is a window into a company's pain. If you see a surge in postings for "Revenue Operations Manager" with mentions of "data silos," that's a signal that the market is struggling with integration. Use sites like In fact or LinkedIn Jobs. For example, I tracked "compliance officer" postings in fintech, they grew 40% year-over-year in 2023 according to Burning Glass Technologies.
  • Funding announcements: Crunchbase and PitchBook track every round. Look for startups building in your space and read their product descriptions. What problem are they solving? Who are they hiring? This tells you where VCs are placing bets.
  • Online communities: Reddit, Quora, and industry-specific forums are goldmines. Search for phrases like "How do I..." or "Is there a tool..." followed by your domain. See how many upvotes the question gets, that's demand in action. I once found a thread on r/sales with 500 upvotes complaining about CRM data entry. That became a product idea.
  • News and press releases: Google Alerts on keywords related to your potential market. Track how often the topic appears in the news. If it's accelerating, you're onto something.
  • Public government data: For regulated industries, SEC filings, patent databases, and state business registries can reveal companies entering or exiting a space.
  • I spend two to three weeks on this before I even think about building. The goal is to create a long list of potential markets, maybe five or six, and then score them based on signal strength. For each market, I assign a score from 1-10 for signal volume (number of job postings, forum posts), signal urgency (how explicit the pain is), and signal growth (trend over time). The highest scoring markets get validated next.

    Step 2: Validating demand with real conversations

    Once I have a list of promising signals, I move to validation. But I don't just call random people. I use the data to find people who are actively searching for a solution. For example, I look at job postings that mention the problem explicitly. If a company is hiring a "Data Quality Manager" and their job description says "we have no central source of truth," that's a qualified lead waiting to talk.

    I reach out with a simple ask: "I'm researching a problem that your company seems to have, [specific pain]. Would you be open to a 15-minute chat?" Because I'm not selling, and because I reference something from their public posting, the response rate is usually over 30%. In those calls, I ask three questions:

  • How are you solving this problem today?
  • How much time/money does it cost you?
  • If you had a magic wand, what would the solution look like?
  • If the answers reveal a painful, expensive workaround that they're unhappy with, I know I'm onto a real market. If they say "it's not that bad" or "we're fine with our spreadsheet," it's a red flag. I once had a call with a VP of Engineering who spent 10 hours a week reconciling data between tools. He said, "If someone solved this, I'd pay $500 a month without thinking." That was the green light.

    But validation isn't just about asking people. You can also test with minimum viable ads. Run a small Facebook or LinkedIn ad targeting job titles that appeared in your public data. Use a lead magnet like a free checklist. If you get a 5% conversion rate, that's strong demand validation. If you get 0.5%, it's weak.

    Step 3: Building a buyer persona from verified data

    After validating demand for a specific problem, I build a data-driven buyer persona. Not the fluffy "Sarah is a 35-year-old marketing manager who likes yoga" nonsense. I mean a persona based on actual attributes from public records:

  • Company size: From LinkedIn or business databases
  • Industry vertical: SIC codes or NAICS from public company registries
  • Job titles of decision-makers: From LinkedIn and job postings
  • Tech stack: BuiltWith or similar services can show what tools a company uses
  • Funding stage: CrunchBase for startups
  • Growth trends: Number of employees, revenue estimates (if available)
  • I combine all this data into a spreadsheet. Then I use tools like ProspectAI (disclosure: I work with them) to enrich the list, automatically appending contact details, social profiles, and company info from public sources. This turns a manual research project into a structured pipeline.

    The key insight: your ICP (ideal customer profile) should emerge from the data, not from your gut. The first time I did this, I discovered that our target market wasn't mid-sized tech companies, it was small law firms. I never would have guessed that in a brainstorming session. Why? Because job postings from small law firms mentioned "compliance tracking" more than any other segment. Public data revealed a niche that our intuition missed.

    And here's a tip: don't just focus on the decision-maker title. Look at the buying center , who else gets involved? Public data from sites like Glassdoor can show you org charts and pain points for different roles. If the "IT Director" is complaining about vendor lock-in but the "Compliance Officer" is worried about audits, you need two different messages.

    Step 4: Testing your positioning before building

    Even after validation, you need to test your messaging. The cheapest way is to use public data for ad targeting. Most social platforms allow you to target by job title, company, and industry. Run a small test with a landing page that describes your solution. See if people click, and more importantly, see if they sign up for a waitlist.

    I recently worked with a client who wanted to sell a compliance tool to fintech startups. Instead of building first, they ran LinkedIn ads targeting "Compliance Officer" at fintech companies with 10-50 employees. They used copy based on pain points from public forum discussions. The click-through rate was 4.5%, ten times higher than their previous campaigns based on generic features. That's the power of data-driven positioning.

    Another test: search interest. Use Google Trends to compare the frequency of different problem statements. Which one is growing faster? If "automated compliance" is flat but "regulatory reporting headache" is spiking, guess which problem to address? You can also use AnswerThePublic to see what questions people are asking about your topic.

    And don't forget competitor analysis using public data. Look at review sites like G2 or Capterra. What do users complain about in existing solutions? Those complaints are your positioning goldmine. If people say "Tool X is too expensive," you can position as affordable. If they say "Tool Y doesn't integrate with Salesforce," make that your differentiator.

    Why public data is the secret weapon for budget-conscious startups

    Let's talk about cost. Premium market research reports can run $5,000 to $20,000. Hiring a consulting firm? Even more. But public data is essentially free. The only investment is your time, and tools like ProspectAI that automate the collection process.

    Consider this: a single job posting that mentions a specific pain point is worth more than a thousand anonymous survey responses. Why? Because it's a real hiring decision, a company is willing to spend money to solve that pain. That's a buying signal. And public data is full of them.

    The startups I've seen succeed with this approach are the ones that treat market research as an ongoing process, not a one-time event. They continuously monitor job postings, news, and community chatter to spot shifts in demand early. They build a feedback loop: data informs outreach, outreach validates data, and the validated data feeds into product and sales.

    You might be thinking, "This sounds like a lot of manual work." And you're right, it can be. But the beauty of public data is that it scales. Once you've identified the right sources, you can automate alerts and even scrape data with simple tools. The key is to start small and expand. A friend of mine built a $1M ARR SaaS by monitoring job postings for "event manager" and noting that everyone struggled with attendee check-in. He built a simple app and launched it in 90 days. Public data gave him the confidence to move fast.

    Frequently Asked Questions

    How do I know if a public data source is reliable?

    Cross-reference multiple sources. If a job posting mentions a problem, verify it with forum discussions and news articles. Government data is usually reliable; user-generated content (like Reddit) should be taken as signals, not facts. For curated insights, pay attention to data reliability from official sources like the Bureau of Labor Statistics.

    What if I can't find any public data in my niche?

    That might mean the market is too small or too early. Consider expanding your definition of the problem. Public data exists for almost everything, sometimes you need to use broader keywords. Also check patent databases for innovation activity.

    Do I need technical skills to mine public data?

    Not necessarily. Most sources have search interfaces. For advanced scraping, you might use a tool like Python or a no-code scraper, but you can start with manually copying data into a spreadsheet. The key is knowing what to search for.

    How does ProspectAI help with this?

    ProspectAI aggregates public data from thousands of sources and provides a unified interface to search for companies and contacts that match your criteria. It's designed specifically for using public data for business discovery, saving hours of manual work. You can enrich leads and build ICPs automatically.

    What's the biggest mistake in this approach?

    Falling in love with your own hypotheses. The data will sometimes tell you that your idea is wrong. Listen to it. The sooner you pivot, the less time and money you waste. Also, avoid confirmation bias, actively look for signals that contradict your assumption.

    Looking ahead: The future of market research

    Public data is only becoming more accessible. Governments are opening more datasets, companies are publishing more APIs, and social platforms are making search more powerful. The startups that will thrive in the next decade are those that build a systematic approach to scanning public data for market signals. It's not about having a single "aha" moment; it's about creating a continuous intelligence stream. So start today. Pick one source, job postings or Reddit, and spend an hour looking for patterns. You might just find your next billion-dollar idea hiding in plain sight.