Tell me how scraping for LLM training observes and respects copyright licensing, including ensuring that no "copyleft" code ends up in proprietary software. Because this is a quagmire.
(I also asked Claude how LLM scrapers respect licensing and it said they mostly don't
)