From 500,000 URLs to Real Jobs: Inside the 80 Days to Stay Engine
An update on the scraper chewing through 500,000 pages overnight, sorting duds from real job leads, and the plan to narrow that pool to 5,000 high-quality companies.
Overnight, a bot ran through roughly 50 sites at a time, in parallel, pulling whatever it could find. By morning it had chewed through a huge slice of an initial pool of about 500,000 pages, and the job now is figuring out what's actually usable in that pile. Some of what came back is dead weight. Some of it is exactly what the project needs.
This is an inside look at the 80 Days to Stay engine, the system behind a job search tool built to turn a mass of scraped web pages into a short list of real opportunities people can actually apply to.
Sorting the duds from the gold
The first pass through the results is basic triage. Some sites return a single page with nothing useful on it, maybe because the content sits behind a login, maybe for reasons that aren't obvious from the outside. Those get marked for a retry later or set aside entirely as not worth pursuing. The next site checked in this pass looks completely different: a real company, clearly identifiable, in this case an insurance company, with a solid about page, detailed company information, and multiple career-related pages, careers one, careers two, a jobs page, specific listings. That's the kind of result the whole pipeline exists to find, a site with enough structure and detail to extract something genuinely useful.
Narrowing 500,000 down to 5,000
The plan is to take that initial pool, roughly 500,000 pages, and narrow it down to around 5,000 high-quality leads. That narrowing isn't random filtering; it means annotating what each company actually offers, what kinds of roles they're posting, including whether a listing labeled "senior" is genuinely senior or just labeled that way, and then pushing all of that structured information into a database. One example found in this pass was a listing for a senior security role. Whether that role is actually out of reach for an entry-level candidate depends on the specific skills required, not just the seniority label attached to the title, which is exactly the kind of nuance that makes manual annotation worth doing instead of relying on the label alone.
Speed now, precision later
The immediate goal is deliberately narrow: get usable lists into people's hands within the next couple of days, not wait for a perfect system. That means taking one pass through everything, getting the initial 500,000 down to a workable 5,000, and shipping that list so people can start applying. The plan is to come back afterward and clean up everything the first pass couldn't get, at a slower, more careful pace.
The longer-term tool: upload a resume, get matched
Past the immediate list, there's a bigger tool being planned. The idea is an intelligent search: someone uploads their resume, and the system scans everything collected in this process to match them against likely opportunities automatically. That's a bigger build, estimated at around a month to get to something sophisticated, which is why it's being treated as the next phase rather than this week's deliverable.
Why build this at all
The reasoning behind the whole project is stated plainly: a tool like this should already exist. Job seekers, and especially those working against a tight timeline, shouldn't have to manually comb through hundreds of company websites looking for real, current openings. Since that tool doesn't exist yet, the 80 Days to Stay engine is being built to fill that gap directly, starting with brute-force scraping and moving toward something closer to automated matching.
Key takeaways
- The scraping bot runs roughly 50 sites in parallel and has pulled data from a pool of about 500,000 URLs.
- Results get sorted quickly: sites needing a login or offering nothing useful are set aside, while sites with clear about pages and multiple job listings are flagged as high-value.
- The plan is to narrow 500,000 pages down to about 5,000 annotated, high-quality leads, tracking what roles each company posts and their actual skill requirements.
- The near-term goal is getting a usable list into people's hands within days; deeper cleanup and refinement come afterward at a slower pace.
- A longer-term tool is planned that lets someone upload a resume and get automatically matched against the collected database, expected to take about a month to build.
Who this is for
This project update is for anyone following the 80 Days to Stay engine, a Humanitarians AI initiative built to help people find real job leads quickly by turning a massive scrape of company websites into a usable, annotated database rather than making job seekers search company by company on their own.
Full transcript(auto-generated, with timestamps)
[0:01]Okay, 80 days to stay updated. So, the bot is still running. It's been running all night. It's 50 in parallel. Uh, but it looks like we have quite a few. These are all Each one is a website that we grab something from. So, let's take a look at a couple of them. This one's just a single page. We look at the page sort of I don't know what that is. So, what I'm going to do with this, maybe it needs a login or something. Who knows? Um, uh, but with this one, I'm just going to mark it like, you know, we need to try again or just, you know, mark this as
[0:40]Not something we care about. If we look at the next one, this looks good. So, we have some about pages. So, clearly, it's an insurance company. I'd be shocked. They've the models that we sent it to. They have career pages. Careers one, careers two, jobs, da da da. It looks like they have specific job listings here. So, this one we can get sort of a lot out of here. And so, uh, my thing is just to sort of take a pass through everything once. It'll take that initial 5,000 500,000 down to maybe 5,000. But I want to get sort of high quality leads to people like this week within
[1:31]The next couple of days. Then we'll go back and sort of clean up everything we can't get. But this company looks, you know, like what we're hoping for. A lot of information about the company. We have career pages. We even have individual job pages. So this is a senior security something. Um so that's probably not an entry level one. Uh but all the skills they're looking for is somebody somebody could do. So it really depends on what they mean by senior. Um, but we'll just annotate all this and then put that annotation in the in the database and then start working on sort of an intelligent search.
[2:20]The search I'm thinking of building is, you know, just upload your resume. It'll go through all this stuff and just start matching you to, you know, what's likely. Uh, but that's going to take a bit. I mean a month to build it to to have a sophisticated tool. So again my goal sort of this week is just to get some list in people's hands that they can work with. After that we can take a slower pace and you know make the tool that should exist. So you know this tool should exist now. I shouldn't have to be building it, but doesn't exist.
More from 80 Days to Stay
4:15Building a AI-Powered Web Scraper to Help International Students Find Visa Sponsors | 80 Days 2 Stay
2:19LLMs Are Like Puppies: They'll Always Fetch, But Sometimes The Wrong Thing | 80 Days to Stay
7:31Why LinkedIn Failed Us: Building a Real Database for International Student Hiring (80 Days to Stay)
10:53Building a Visa Sponsorship Database from SEC Data | 80 Days to Stay Project
14:1880 Days to Stay
2:08