That's not going to save them, the model could leak GPL licensed code on output as well, and it's an order of magnitude more likely than leaking internal documents.
These large tech companies can't have their cake and eat it too.
It's not new Anthropic/OpenAI/whatever trained on copyrighted code and leaks (anna's archive). Companies just don't give a shit because you can't prove it.
In Asahi Linux specifically, there was not long ago a PR for M3 or similar for removing a big roadblock that existed (and I think still exists). Turns out this whole PR was implemented by an LLM (using copyrighted materials, probably most of it is Apple's) and it was immediately closed. Yes, it's a problem and illegal.
If you can't prove it, you can't prove it both ways.
Either everybody gets to use it or nobody does.
Especially that the code getting trained on is for the vast majority copyrighted open source software so if anybody has something to say, it's the opensource community rather than Apple
These large tech companies can't have their cake and eat it too.