Microsoft’s Copilot chatbot has cleared a significant hurdle in its copyright row with The New York Times, claiming that its AI rarely reproduces more than 16-word chunks from the newspaper’s articles. In a trove of 8.2 million chat logs, only 0.0074% contained at least 16 words in common with NYT content, the company asserts.
As part of the lawsuit, Microsoft shared these logs with experts, who analyzed them for evidence of copyright infringement. The findings, however, paint a picture of minimal overlap, with just 59,545 logs containing any shared text, and only 24 instances containing at least 30 matching words. This, Microsoft argues, supports its case for fair use of copyrighted material for AI training.
Despite these findings, The New York Times is unconvinced. Ian Crosby, the Times’ lead counsel, stated that the documents and testimony reveal Microsoft and OpenAI’s use of the newspaper’s content ‘to make commercial products that substitute for its journalism, threaten its business, and undermine its industry.’
Microsoft’s legal position is that while Copilot relies on copyrighted material, the resulting systems serve different purposes. The company argues that the occasional text replication does not detract from the transformative nature of AI training.
The legal battle is ongoing, with Microsoft seeking summary judgment, which could potentially end the case early. Meanwhile, the conversation around AI and copyright continues to evolve, with implications for the future of journalism and technology.







