RESEARCHInvestigateNOW
One prompt is not enough: Instruction Sensitivity Undermines Embedding Model Evaluation
arXiv cs.CL — Computation and Language
Factual evidence
What the source reports
Study shows instruction embedding models are highly sensitive to prompt phrasing, making single-prompt evaluations unreliable across tasks.
Open source