3 comments

  • infotainment 52 minutes ago
    The "distillation" argument always seemed pretty weak to me. If one argues that training an AI on copyrighted content is merely analogous to reading a book (and therefore is fully transformative), then it stands to reason that training an AI on other AI outputs would also be fully transformative.

    It seems to me you can't have it both ways; either training is a violation of copyright or it isn't, and there's no consistent argument where "distillation" is a violation but other things aren't.

    • easterncalculus 16 minutes ago
      This argument only makes sense if you equate a printed book with a chatbot. They aren't the same thing, and the chatbot's outputs aren't copyrighted.
  • andsoitis 41 minutes ago
    How does distillation actually work? If I wanted to distill a model and I have enough money and compute, what’s the process?
  • jqpabc123 56 minutes ago
    So China can't steal the data that US companies stole fair and square?

    https://www.techopedia.com/trump-administration-openai-train...