Member of Technical Staff (Data Scientist, Evals)

Perplexity · Remote

Posted
55 days ago
Last confirmed live
1 day ago

What this role involves

This role involves building automated evaluation pipelines to assess answer quality across Perplexity's products. The data scientist will design evaluation sets for tool calls and develop VLM-based solutions for visual rendering evaluation. They will work in a small team to define metrics and incorporate public benchmarks.

Skills this posting asks for

  • python
  • sql
  • aws
  • databricks
  • llm as judge
  • vlm
  • evaluation metrics
  • factual consistency
  • hallucination rate
  • retrieval precision
  • ground truth datasets
  • agentic coding
  • ai-assisted development tools
  • research methods
  • machine learning

Requirements

  • 4 years of experience
  • Level: mid

From the employer’s posting

Perplexity serves tens of millions of users daily with reliable, high-quality answers grounded in an LLM-first search engine and our specialized data sources. We aim to use the latest models as they are released, but the intelligence frontier is a jagged one, and popular benchmarks do not effective…

Read the full description on Perplexity’s careers page

Apply without filling the form

Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.

Other roles at Perplexity

All 20 roles at Perplexity