The paper argues that when language models are created using…
The paper argues that when language models are created using training data collected without adequate documentation (data sheets) about its composition, potential biases, and filtering methods, this leads to an ethical risk they call “Data Amplification,” where errors are replicated across the web.