Building a Visa Sponsorship Database from SEC Data | 80 Days to Stay Project

This video builds a company database from quarterly SEC Form D filings, then shows how LinkedIn hiring patterns can flag which funded companies actually sponsor visas.

10:53 video4 min readWatch on YouTube

International students face a real employment gap, and part of the reason is invisible rather than structural: the funded companies most likely to actually sponsor a visa are largely absent from the job boards students already search. This video is day two of the 80 Days to Stay project, and it walks through the first real technical step: pulling raw SEC filing data by quarter, cleaning it up, and starting to figure out which of those companies have a track record of hiring internationally.

Pulling filings quarter by quarter

SEC filings are organized by quarter, so the script built here works by iterating through each quarter going back roughly ten years, far enough to capture companies with an established funding history without reaching all the way back to when SEC filing data begins. For each quarter, the script pulls how much money a company raised in that filing period. Because a company can raise money in multiple quarters, for example once in Q1 and again in Q3 of the same year, the script updates that company's totals rather than treating each filing as a separate entity, removing duplicates and adding up funding across filings as it goes. At this stage nothing gets filtered out; the goal is to keep everything, across every industry and every state, and let the data volume settle before deciding what to cut.

Descriptive statistics as a first checkpoint

Once the pull finishes, the script outputs descriptive statistics on the result: total company counts, funding by state, and breakdowns like how many companies raised above $1 million or above $5 million. This run turned up a little over 500,000 companies, with about 250,000 of them having raised over $5 million. A meaningful chunk of that total, flagged under an SEC filing code referred to as "E9," corresponds to companies based in Europe, which get set aside since the project's focus is US-based employers. Looking at concentration by state lines up with expectations: California, Massachusetts, New York, and Texas show up as major hubs, alongside states like Texas and Illinois more broadly.

Narrowing by industry

A parallel breakdown looks at industries, aiming to identify the sectors most likely to actually hire recent graduates and international students specifically. Pharmaceuticals, biotech, and computers stand out as strong likely candidates, while others need more digging, investment funds are assumed to be finance-related and might hire analysts and programmers, while real estate looks like a weaker fit. Adding a handful of these promising industries together lands on a working set of roughly 50,000 companies as a realistic starting point for the next stage of filtering.

Checking for a real hiring signal via LinkedIn

Counting funded companies in the right states and industries only gets you a list of plausible candidates, not evidence that a company actually hires internationally. The next step described is a manual, illustrative pass through LinkedIn data at a real, already-identified startup, looking at employee profiles for a specific pattern: a US-based master's degree paired with an undergraduate degree earned overseas, plus signals like language skills listed on a profile. That combination is treated as a reasonably strong indicator that a person came to the US in part for graduate study, which in turn means that if a company has hired someone matching that pattern, it has functionally demonstrated willingness to hire international candidates, since it has already done so. The same logic extends to recent-graduate hiring: if a profile shows a bachelor's degree from just a couple of years ago, that's evidence the company is open to hiring people early in their careers too.

Where this is heading

The plan is to automate this LinkedIn-pattern check across the full filtered company list, likely pulling company websites and public profile data to look for the same degree and language signals at scale, in order to build a predictive score for how likely a company is to sponsor a visa or hire someone with a similar academic background. The mission for the full 80 Days to Stay project is to facilitate at least one successful sponsorship match using only public data and open-source tools, on a stated budget of zero dollars per month.

Key takeaways

  • SEC Form D filings are organized by quarter, so the data pipeline pulls, deduplicates, and aggregates funding totals quarter by quarter across roughly a ten-year window.
  • An initial, unfiltered pull turned up a little over 500,000 companies, with about 250,000 having raised over $5 million; European filings (flagged under a code called "E9") are excluded from the US-focused target list.
  • Industry filtering narrows the target list toward sectors most likely to hire recent graduates and international students, landing near 50,000 companies as a starting point.
  • A manual LinkedIn review technique looks for employees with a US master's degree and an overseas undergraduate degree as a signal that a company has already hired internationally.
  • The next planned step is automating that LinkedIn-pattern check at scale to build a predictive score for sponsorship likelihood across the full company list.

Who this is for

This video is part of the 80 Days to Stay project from Humanitarians AI, aimed at international students facing OPT deadlines who need a more targeted way to find companies actually capable of and open to visa sponsorship, beyond the well-known names everyone already applies to. It's also a useful watch for anyone interested in building a public-data pipeline from scratch using SEC filings and open-source tools on a zero-dollar budget.

Full transcript(auto-generated, with timestamps)

[0:01]Bear here. It's 80 days to stay. So what you're looking at here is a little script. Uh what the script does is the SEC is organized by quarter. So if you get money in a particular quarter, you have to file that quarter. So what this does is the data is all by quarter. So we just run a little script which guess you know quarter uh fourth quarter 2004, third quarter 2004. We went back, I think, to 2015. We could go back to 1970 or 50 or whenever it starts, but we won't. We'll go back 10 or so years because we don't really care about startups. We care about companies that

[0:41]Have money. And so, what this does is it goes through there, it looks and it, you know, says, "Well, how much money did you raise?" And then updates values because sometimes you raise quarter 1, raise more money quarter three. and then you have to file both quarters, but it's the same company. So, this is going through, it's removing duplicates, it's adding things up, it's getting statistics. We're not filtering at this point. So, it's getting a lot of companies, you know, uh 11,000, 9,000, 8,000, 9,000, etc. But a lot of those are very small company. So, but we do have we're just keeping all the data at this point

[1:22]And then we'll filter it down. But because we don't know exactly our target. Do we want to filter to a million? Do we fil to 5 million? What is our minimum threshold of a company that's likely to hire? But we're just going to grab everything. We're grabbing everything in every industry in every state for the past 10 or so years. And so uh this will run my guess is the next 10 minutes we'll have I don't know um because there are a lot of the same company raising money again and again so they'll appear again and again. It'll also output descriptive statistics how much money you know what states have

[2:05]What how many you know what's the average how many above 5 million how many above a million or whatever. So, day two, uh, 80 days to say, it's gone good so far. I'm going to talk tomorrow in more detail about how are we going to map this to companies that are likely to hire international people. We already have that planned out. Uh, but, uh, it's laid down to Friday, so I probably won't talk about that right now. talk about that tomorrow, day three of which we'll filter it down to maybe a hundred thousand or a couple hundred thousand which have enough revenue to actually hire people

[2:42]And then we'll go from there. Um I think it'll be fairly straightforward the two other things we want and that's have they or are they likely to hire international or and are they likely or have they hired recent graduates but I'll talk about how to do that tomorrow um because it's Friday late and uh it's been a long day. Okay, here are some descriptive statistics of what we generated today. uh we found a little over 500,000 companies about 250,000 of them or over 5 million from from Europe. That's what E9 means. We're going to eliminate these. Um obviously at our st this E9 thing we're about Texas, Massachusetts. So do

[3:27]California, New York. That's sort of what we expect. We'll have to look at the industries. I did another breakdown of the industries here. Um, so we're going to target specifically on on sort of industries which are likely to hire recent graduates and international students. I'll have to dig into some of these industry like an investment fund. I'm assuming that's like in finance. Uh, we'll have to see whether they're hiring, but maybe because they're likely to hire analyst and programmers and people like me. Real estate, I'm not sure. probably low. Um, but pharmaceuticals for sure, biotech for sure, computers for sure, finance for sure. So, we'll

[4:14]See how many companies we end up with. But, you know, 50,000 companies uh if we add a couple more industries here is sort of our starting spot. So the idea is to, you know, start with these 50,000 companies and start to really annotate whether or not uh they're likely to to uh sponsor and have hired recently. I have to look into the LinkedIn data a bit more. Let me show you sort of a rough approach. So this is a very smart guy who was a former student of mine who's working for a startup. So if we look at the startup is something called LEO. So if we look at the startup if we get

[5:03]Information about the people we can make sort of emphasis of whether they hire. The way we do that is we look at their employers. So if we look at somebody like this guy so he has a masters in computer science. He he um graduated when did he graduate? graduate in 2020. So he is somewhat recent but very often he doesn't. So he says he speaks Hindi fluently. That's sort of the indication that you're international. Another thing let's take a look at Dev's resume directly. So what we'll look for is patterns in these these. So for example, Ed Dav, he's a more typical thing. So he got a US US degree here.

[6:00]But if you look at his undergraduate, it's from India. So this pattern of getting a US master's and having a undergraduate from a different c uh country is a pretty in good in case you're that you're on a visa and doing a visa thing. And so we'll look at these patterns. And so this essentially tells us that okay, this group is open to hiring international students because they just hired one. And so what we're going to do is we're going to automate this. So what we're going to do, I'll have to look into the LinkedIn API or other ways to get this data other sources of data like this.

[6:39]But the idea is basically once we have a company name to find other information like their website, their their company page, dot dot dot dot, and then once we have their company page here, we start looking at the people. This looks like a pretty mint company. I'm just looking at the names, but what a bot can do is they can dig in deeper. So my guess is this is not an international guy, but I could be wrong. What we do is the Bible come and dig further and and so and good thing about these things is they often give education. So this person did his undergraduate in

[7:18]America and his graduate in America. So there's a good chance that he's American and the name and everything else um indicates that you know he's probably American. So we'll develop a a predictive model for that. I'll talk about the details of how we do that. But a lot of it is logic. So you could also not develop a formal model but just a little bit more logic. Basically take a look at somebody else. Uh this person here. So if we look at this person here, they got a masters in America and an undergraduate overseas. They're very likely to be international. That combined with the name combined

[8:04]With the languages they likely speak they list them English and karaji they're America's lesbian karach particularly fluent quote professional. So I I uh he doesn't say he's fluent. My guess is he's probably as fluent. Um uh but we'll we'll do those things. And so what that'll allow us to do is if we see like this company uh if our model is doing its job should give us a very high score uh of its willingness to hire international students. If we see a different company and it has a different you know look if it looks like they've never hired anyone then it would have a love score.

[8:45]Uh the other thing we want to indicate is are they willing to hire people who recently graduated and again if we look at a a if we can get this kind of data from LinkedIn it's fairly straightforward because this guy does not list his oh he did so it's that he got his uh bachelor's in 2023. So that means they've hired somebody who got his bachelors a couple years ago. So that basically it states that they're willing to o open to hire recent graduates because they did. They hired a recent graduate. So that's that's the next step. Our next step is to take all of

[9:26]This data here, eliminate the stuff that's very unlikely just from the fact that's in Europe and doesn't really apply or it's in an industry that sort of doesn't traditionally hire tech people and really to focus on those industries that do hire tech people. get the company names and then go to the company websites, company Twitters, company everything and start looking well who's in that company. If it looks like, you know, people who, you know, got a bachelor's degrees overseas and then got a masters in the US, which is very very typical of of somebody, you know, coming to the US for their masters to work in the US.

[10:06]Uh then they give the strong evidence that they're willing to hire because they have hired and so then that you know becomes a target for somebody who's looking for a job. Well, here's a company that's already hired. They they have a couple students who have you know similar background to yours and they've already hired somebody like you at least from your background from that perspective. And then um then that gives somebody a you know a likely target somebody who seems at least willing to entertain somebody who needs you know uh to deal with the visa stuff. Okay, that's that. Take care. Uh we'll pick it

[10:50]Up for day three tomorrow.

More from 80 Days to Stay

Humanitarians AI Lyrical Literacy Project