Multimodal ModelsEvaluation & RobustnessAI Safety

I build and evaluate language and multimodal models, with a focus on reasoning, robustness, and safety.

Open to work

I'm looking for a full-time Research Scientist role in industry, working on AI safety, evaluations, and multimodal models. Based in the UK · open to UK-based or remote roles · get in touch

About

I'm a Postdoctoral Researcher at the University of Edinburgh, working with Pasquale Minervini on long-context language and vision-language models, with an emphasis on evaluation, robustness, reasoning, and AI safety. I completed my PhD in NLP at Edinburgh (EdinburghNLP, CDT NLP), advised by Frank Keller and Hao Tang.

Most recently I was a Research Fellow at Anthropic, studying model behaviour and dual-use risks in frontier LLMs: how seemingly benign generations can be repurposed for harm, and what that means for evaluation and deployment safeguards. I also co-mentor a SPAR project on multimodal safety and mechanistic interpretability in encoder-free vision-language models.

Previously, I was an Applied Science intern at AWS AI, working on hallucination mitigation in LLMs. Outside research, I enjoy photography.

Research Work in progress →

01

Multimodal models

How vision-language models understand images, video and documents, and where they fall short, from reading clocks to summarizing scientific posters.

02

Evaluation & robustness

Benchmarks that expose where vision-language models break: corruptions, long context, and inverse scaling.

03

AI safety & misuse

Dual-use risks in frontier LLMs, and safety evaluation of LLM agents.

04

Long-context & reasoning

Long-document understanding with memory-efficient end-to-end training, and whether chain-of-thought reflects what actually drives a model's answer.

In the media

Lost in Time, our study of clock and calendar understanding in multimodal LLMs, was cited in Stanford HAI's 2026 AI Index Report and covered internationally.

News

  • 2026Do Composed Image Retrieval Benchmarks Require Multimodal Composition? accepted at NeurIPS 2026 Datasets & Benchmarks.
  • 2026VLM-RobustBench accepted at ICML 2026.
  • 2026Lost in Time cited in Stanford HAI's 2026 AI Index Report.
  • May 2026Completed the Anthropic Fellows program.
  • Jan 2026Completed my PhD at the University of Edinburgh.
  • Nov 2025Joined Anthropic as a Research Fellow in London.
  • Sep 2025MMLongBench accepted at NeurIPS 2025 Datasets & Benchmarks Spotlight, plus one NeurIPS workshop paper.
  • 2025PosterSum accepted at AACL 2025.

Selected publications Scholar ↗

All publications
Patents
  1. US10599864B2: Sensitive audio zone rearrangement for customer verification.
  2. US10296523B2: System and method for estimating temporal importance of data.
  3. US10198322B2: Method and system for efficient selective backup strategy in an enterprise.
  4. US20150381703A1: Automating a process for web-based software.
  5. US20160269417A1: Dynamic data masking for mainframe application.
  6. 201621003887 (IN): Systems and methods for estimating skill-sets of users in a distributed environment.

Contact

Reach me at

rohitsaxena.uoe@gmail.com

Edinburgh, UK · CV on request

© 2026 Rohit Saxena