RESEARCHInvestigateNEXT 12 MONTHS
Provable Training Data Identification for Large Language Models
arXiv cs.LG — Machine Learning
Factual evidence
What the source reports
Researchers propose a provable set-level inference framework to statistically identify whether specific datasets were used to train an LLM.
Open source