Description
About Us
Sieve is the only AI research lab exclusively focused on video data. We combine exabyte-scale video infrastructure, novel video understanding techniques, and dozens of data sources to develop datasets that push the frontier of video modeling. Video makes up 80% of internet traffic and has become the enabling digital medium powering creativity, communication, gaming, AR/VR, and robotics. Sieve exists to solve the biggest bottleneck in growth of these applications: high-quality training data.
We've partnered with top AI labs and did $XXM last quarter alone, as a team of just 15 people. We also raised our Series A last year from Tier 1 firms such as Matrix Partners, Swift Ventures, Y Combinator, and AI Grant.
About the Role
As an applied research engineering intern at Sieve, you’ll help build high performance building blocks and large scale pipelines to understand video with high precision at internet scale. Often this involves working on ambiguous research problems and finding clever techniques to solve them. You will be working in the computer vision, audio processing, and text processing domains.
You’re likely a good fit if you’re comfortable working with models + APIs and squeezing every drop of performance out of them through clever pre/post-processing, parallelism, pipelining, inference optimization, and occasionally fine-tuning.
Requirements
Excellent communication skills
Deep passion for the video domain and media technologies
Motivated by building end-to-end products—not just training models
Able to break problems down from customer level impact to necessary building blocks.
Bonus: Active contributor to open source projects
In-person at our SF HQ
Benefits
Breakfast, Lunch, and Dinner covered and your choice of snacks
Ubers covered home
Team outings & off-sites!
Company
Sieve is a video data research lab that curates and licenses large-scale video datasets for AI training. It records video from scratch and aggregates it from multiple sources, filters for quality, indexes billions of videos with detectors and embeddings, and annotates with dense labels and pairings. Customers include leading AI labs, Fortune 100 companies, and fast-growing AI startups. The company provides ready-to-use datasets or custom datasets, free data samples, and purchase-based access with SLA-based delivery via S3-compatible transfer, emphasizing compliance and scalable, secure delivery.
Related postings
Hex
Product Engineer InternSan Francisco, CA, USASieve
Software Engineering InternSan Francisco, CA, USACluely
Engineering InternSan Francisco, CA, USABillionToOne
AI Engineering InternUnited States