Description
About Us
Sieve is the only AI research lab exclusively focused on video data. We combine exabyte-scale video infrastructure, novel video understanding techniques, and dozens of data sources to develop datasets that push the frontier of video modeling. Video makes up 80% of internet traffic and has become the enabling digital medium powering creativity, communication, gaming, AR/VR, and robotics. Sieve exists to solve the biggest bottleneck in growth of these applications: high-quality training data.
We've partnered with top AI labs and did $XXM last quarter alone, as a team of just 15 people. We also raised our Series A last year from Tier 1 firms such as Matrix Partners, Swift Ventures, Y Combinator, and AI Grant.
About the Role
As a distributed systems engineer at Sieve, you’ll design and engineer systems that handle the compute, scheduling, and orchestration of complex ML + ETL pipelines that need to run quickly, reliably, and cost-effectively on large sums of video.
You’re likely a good fit if you love optimizing for system uptime, have worked with cloud technologies, optimizing hyper-fast distributed systems at the scale of thousands of GPUs, and building great internal tooling and CI/CD for rapid iteration.
Requirements
3+ years of experience building foundational data infrastructure
Proficient in working across diverse cloud architectures
Designed and maintained pipelines that process petabytes of data
Developed robust CI/CD pipelines tailored for ML-focused teams
Strong coding experience with Go and Python; Experience with Rust is a plus
Operates as an IC who leads by example
Experience with large-scale video data systems
In-person at our SF HQ
Benefits
401k + Full Health Insurance
Breakfast, Lunch, and Dinner covered and your choice of snacks
Ubers covered home
Company
Sieve is a video data research lab that curates and licenses large-scale video datasets for AI training. It records video from scratch and aggregates it from multiple sources, filters for quality, indexes billions of videos with detectors and embeddings, and annotates with dense labels and pairings. Customers include leading AI labs, Fortune 100 companies, and fast-growing AI startups. The company provides ready-to-use datasets or custom datasets, free data samples, and purchase-based access with SLA-based delivery via S3-compatible transfer, emphasizing compliance and scalable, secure delivery.
Related postings
Palo Alto Networks
District Systems EngineerSan Francisco, CA, USADoorDash
Systems EngineerSan Francisco, CA, USA and 1 otherAnthropic
IT Systems EngineerNew York, NY, USA and 2 othersCloudflare
Software Engineer: Distributed Systems (Infrastructure)Austin, TX, USA and 1 other