Basically, I started to search myself. I wanted to buy my own business. What I found was that all these various scrapers out there couldn't track my local brokerages, like the Sacramento business brokers and the small M&A outfits. So I built scrapers to do that.
That's turned into the thing I've been chasing ever since. I go after what I call the outpost brokers. These are broker websites that have like 3 or 4 listings, and they only get 1 new listing a quarter. I catch that listing when it happens. And because I also scrape all the big aggregators, I can tell you when something truly is an outpost listing and it's going to be less competitive.
That's alpha, at least in theory. I wanted to know whether it shows up in the numbers.
So a third party pulled 21 days of deals off a well-known deal-sourcing platform, filtered to EBITDA of $1.5M or more, 200 rows, and sent me the CSV. I ran the same window and the same threshold against our database. The file was de-identified, no broker, no URL, no source, so matching had to happen on financials, location and timing. I went through all 200 by hand.
We had 163 of their 200 rows.
We also had 62 qualifying deals that weren't in their file at all. Those 62 came from 19 different sources, and no single source accounted for more than 26% of them. On their side, 69% of the deals traced back to one big marketplace.
That's the whole bet, sitting in one number. Breadth is a lot of small pipes, and it's tedious, and it doesn't compress into a single integration.
Now the part that didn't go my way. They had roughly 16 to 19 deals I didn't have. Most of them had no city and no state, just "Mid-West" or "Multiple Locations." That reads like off-market sell-side advisor inventory, the kind that never touches a public broker page. I don't have that today and I'm not going to pretend otherwise.
A few other things fell out of reviewing 200 rows one at a time. Their 200 collapsed to about 188 distinct deals. One surgical center portfolio appeared 4 times. 12 rows were already marked Closed, Sold, or Under LOI, so they were still sitting in a feed I was supposedly competing with.
Then the estimates. 142 of their 200 EBITDA figures, 71%, were flagged as estimated by their own platform. Where they published a real reported EBITDA, it matched us on all 31 I could check, so their real data is accurate. Where the figure was estimated, the median came in at about 2.2 times what the listing itself reported.
Let me tell you a little bit of inside baseball here, because I estimate missing fields too, and I do it on purpose.
A lot of people run a search based on cash flow. If you do a cash flow search, you're missing every deal that doesn't have cash flow in it. That's stupid. A lot of people search on price, but some sellers won't publish a price, so you miss that deal too. So here's how I solved the problem. I looked at industry benchmarks, and if you give me price, cash flow, or revenue, the AI estimates the missing pieces so the deal shows up in your feed regardless. The number might be totally raw, but if you have a wide enough band, it's going to show up on your list.
Estimating so a deal stays visible is good. The failure is when an estimate gets handed to you as if it were a reported number, because then it silently clears a filter it never should have cleared. You set your floor at $1.5M and you get deals the listing itself never claimed.
One last thing, and it's a real ask. I take requests on which websites to add. Customers send me sites all the time and I put them in the queue.
So if you know a small regional brokerage, or some niche industry site that only ever lists HVAC shops or veterinary practices, send it to support@searcheros.ai and I'll go look. I probably don't have it. The outposts are the whole reason this works and there's no list of them anywhere, so I find them one tip at a time.