Mountain View, California, USA
2 days ago
Research Intern - Training-Time Provenance (Data Dignity)

Research Internships at Microsoft provide a dynamic environment for research careers with a network of world-class research labs led by globally-recognized scientists and engineers, who pursue innovation in a range of scientific and technical disciplines to help solve complex challenges in diverse fields, including computing, healthcare, economics, and the environment.

Training-time provenance is a research effort on estimating the influence of specific training data on outputs of large language models (LLMs). Current neural network architectures are opaque in terms of providing sources for their generations, and there are at least two good reasons to change this:

“X-ray” into intent, so that we can detect bad human actors or dangerous AI activity by identifying the most influential source documents related to a given model output. For instance, sneaky prompts might invoke articles about bomb making that could evade guardrails otherwise. This will be a deeper method of countering this type of danger than others currently in use.“Data dignity”, meaning incentives, recognition, and potentially pay for people who contribute certain valuable data to unforeseen kinds of models we will want in the future, assuming the future will surprise us fundamentally. The goal is to foster new classes of creative professionals where possible, instead of relying solely on ideas like Universal Basic Income in the event of a future with very high-functioning large models. 

We are attempting to demonstrate that LLMs can be trained in such a way that influence of specific training data on generated outputs can be efficiently and usefully estimated. You can read more about “Data dignity” in the article: There is no A.I. (The New Yorker).

Confirm your E-mail: Send Email