Skip to jobs
Unlisted
All jobs

prometheus · R&D

Member of Technical Staff - Multimodal Pre-training

Where
San FranciscoHybrid
Experience
8+ yrsguessed from the title
Pay
Not stated
Posted
Already listed when Unlisted started watching this board (6 Oct 2026)
Apply on AshbyOpens the employer's own posting

Checking Ashby for this posting…

About the Role

This role owns how non-text modalities enter the pretraining run: the data we train on, the encoders and fusion architecture that carry it into the language model, and the capabilities we get out.

You will build and scale multimodal data pipelines (images, video, 3D, physical / scientific data), run the architecture research that decides how modalities are tokenized, encoded, and interleaved with text, and define the evaluations and ablations for validating your data and architecture. The work spans the full stack, from a data or architecture hypothesis to a controlled training experiment to a verdict that lands in the flagship recipe.

What You'll Do

What We're Looking For

Why This Role Matters

Multimodal capability is one of the clearest frontiers left in model quality, and this role controls the full path from raw data to shipped capability. The decisions made here determine what the model can see, watch, and understand beyond text.