Company
Wynd Labs
Annual salary
$160k – $250k/yr
Location
Remote
Listed
121d ago
- Experience:
- 2+ years
- Workplace:
- Fully remote team with global hiring capabilities.
- Visa:
- Visa sponsorship not available
- Equity:
- Competitive equity
Remote. Visa sponsorship not available.
Send us your LinkedIn and CV. If your experience fits, we'll introduce you to the recruiter filling this role.
Apply for a referral →The recruiter emails you before anything happens and may suggest other jobs that suit you better.
What you'll do
- Build and maintain large-scale web crawlers across diverse domains (social media, travel, multi-language sites) that power dataset creation for frontier AI labs
- Design high-throughput, fault-tolerant systems for data collection handling millions to billions of URLs/day
- Handle anti-bot systems, rate limits, and dynamic/JS-heavy sites, thinking creatively when standard protocols fail
- Develop pipelines for cleaning, deduplication, filtering, and normalization of web data at TB, PB scale
- Construct and maintain datasets for research and model training, collaborating directly with research teams to align data collection with modeling needs
- Monitor crawl performance, coverage, and data quality; iterate quickly as web environments constantly change
- Optimize infrastructure for cost, latency, and reliability across cloud or bare-metal environments
What you need
- We need someone with 4+ years of experience (2+ for candidates with PhDs) building web crawlers and large-scale data acquisition systems who has hands-on experience designing high-throughput, fault-tolerant pipelines. You should be comfortable operating at the boundary of scale and reliability in adversarial web environments and have a track record of processing millions to billions of URLs/day. Bonus points if you have experience with NLP pipelines, LLM pretraining data, or dataset curation for ML.
About the team:
We build infrastructure that delivers massive amounts of web data to the companies training the world's most powerful AI models. Frontier AI labs are our customers, you'll be working as an extension of their data teams on pre-training and inference model development. We own one of the largest repositories of public web data and have more resources available for working with data at scale than basically any other company. We're a lean, flat organization (~42 people) with no people managers, just builders pushing to expand what's possible for open web data and AI. We're cash flow positive and growing quickly.
Frequently Asked Questions
How do I apply for the Research Crawling Engineer (Remote) role at Wynd Labs? +
Use the Apply for a referral button on this page to send us your LinkedIn and CV. If your experience fits, we'll introduce you to the recruiter filling this role. They'll email you to check you're interested, then put you forward for this job or others that suit you better. It's free.
What does this Wynd Labs role pay? +
The listing gives $160k – $250k/yr.
Is this role remote? +
Yes, the listing is remote. Fully remote team with global hiring capabilities. Visa sponsorship not available.
Related Roles
Browse all startup roles
Salaried roles at startups, AI companies and established businesses, all filled by referral. One application covers every role.
View all startup roles →