Filtering Millions of H-1B Records to Find Real Hiring Signals
A look at why H-1B visa sponsorship records from federal databases reveal which companies actually hire, and how to filter six million-plus filings into something usable.
Funding databases tell you which companies have money. They don't tell you which companies are actually hiring. That gap is exactly what a new data source closes: if a company files for H-1B sponsorship, that filing goes straight into a government database, and it turns out to be a far better hiring signal than who raised a round.
Why sponsorship beats funding as a signal
The insight here came from a simple observation: when a company sponsors an H-1B visa, it has to go through a formal government process, and that process creates a public paper trail. Unlike startup and funding databases, which only tell you who has capital, sponsorship records tell you who has actually gone through the work of hiring someone and backing that hire with a federal filing. That makes it a higher-quality signal for identifying companies that are genuinely bringing people on, not just companies that raised money and might expand someday.
The scale of the data
These aren't small datasets. Beyond the startup and funding databases already being used, there are records from the Department of Labor, USCIS, and other federal sources covering every employer that has ever submitted an H-1B application. The numbers run into the millions, with sponsorship history data that looks to be in the range of six million-plus entries. That scale is exactly why the data can't be used raw. It needs to be filtered down to something a hiring-signal pipeline can actually work with.
What to filter for
Pulling in millions of records only helps if they can be narrowed to what matters. The plan is to check each filing for several things: the company's sponsorship history, the filing dates, the frequency of filings (a company that has filed once looks very different from one that has filed five hundred times), whether the company is still active, and how recent its most recent sponsorship activity is. Some of the companies in these datasets will be out of business by now, but since the records include filing dates, it becomes possible to filter down to recent, relevant activity rather than treating every historical filing as equally useful.
Where this fits in the pipeline
This data source is set to slot into the pipeline after the data is pulled in, joining the other sources already being collected: scraped websites, job pages, social links, company addresses, and more. It becomes one more high-quality layer in a broader system built to surface real hiring signals rather than guesses based on funding alone.
Key takeaways
- H-1B sponsorship filings are public government records, since sponsoring an employee legally requires going through a formal filing process.
- Federal sources like the Department of Labor and USCIS hold sponsorship data for every employer that has ever filed, running into the millions of records.
- Sponsorship history is a stronger hiring signal than funding data because it reflects actual hiring activity, not just available capital.
- Filtering criteria include filing frequency, recency, and whether the company is still active, since not every historical filer is still operating.
- This data is meant to be combined with existing sources like scraped websites, job pages, and social links, not used as a standalone signal.
Who this is for
This is for anyone building lead generation or hiring-signal pipelines who wants a look at how public government data can supplement funding and scraping-based sources. It's a behind-the-scenes look at the ongoing 80 Days to Stay data project.
Full transcript(auto-generated, with timestamps)
[0:00]Tag, the scraping is going well. Um, that should get done over the next couple days. Tomorrow's Thanksgiving, so I'm not sure if I'm going to do anything tomorrow. Um but Sarab who does great things that humanitarians mention what should have been obvious from the very beginning uh but he pointed out that when somebody actually f files an H1B sponsorship that of course goes into a government database and so besides these startup databases funding databases there are databases of the department of layer, labor, HSC, CIS, etc. for companies with the company info of who actually, for example, submitted H1B applications. So, we'll add this data as well. Um, I
[0:58]Thought I got some stats here. I think there's some stats down here. Um, but this can include a lot of companies. Um, I thought I asked. Oh, here we go. Um, they're not the stats here, but I I think this might be like sponsorship history might be like in the million or millions. Um, I thought I saw that. But whatever it is is whatever it is. Um, but this is a nice nice database because it's not just who has money, which is what the other one is about. It's about who sponsored in the past. And this is government debt. So if you file for, for example, H1
[1:47]Visa, you have to go through the government process to do it. And so we're going to grab these data as well. So this this is government data. We'll take a look at it in more detail. when we start looking at these data sets, but we're basically just going to grab whatever information we can from data sets which are requirements of the government. If you're sponsoring somebody, for example, H-1B, you have to file forms and that links you back to the the company. So, does it everyone that's ever filed? So, this is why this is the millions. will probably filter that down to I'm assuming that has dates like who is
[2:28]Filed in the last couple years. Uh some sort of sponsorship thing. Oh yeah, 6 million plus records here. But again, some of them may be companies that are out of business. Um but I'm sure they have dates like a date of when it was filed. And so we can just filter that down to and just sort of the frequency. you know, this company has filed one or they filed 500. So, we'll start looking at that um probably after Thanksgiving is begin to bring in these these data sources as well. Okay, that's it. Take care.
More videos
2:08Bridging the Pixel Gap in Browser Automation.
2:23How One Narrow Safety Rule Can Make an AI Less Safe Everywhere Else.
2:04Why splitting a chunk from its document makes it retrieve for the wrong question
4:20Three You Can Take Back. One You Can't.
2:21Why a 50-turn agent pays for the same screenshot 35 times unless it caches the pixels
1:53