Safety of LLM Agent Harnesses
Overview
Evaluating the safety of language models deployed inside agent harnesses.
This work is in progress. Details and results will be shared here once it is published. Get in touch if you'd like to discuss it.
Results
Results will be posted here once the work is published.
Updates
- Sep 2026Project started.
Cite
Work in progress. If you refer to this project before a paper is available, please cite this page:
@misc{saxena2026harnesssafety,
title = {Safety of {LLM} Agent Harnesses},
author = {Saxena, Rohit},
year = {2026},
howpublished = {\url{https://saxenarohit.github.io/projects/harness-safety/}},
note = {Work in progress}
}
© 2026 Rohit Saxena · Questions about this project?