Copyright in model-generated code is unsettled — what can a team still decide for itself?
answer
- Two questions, only one of them open
- Read what you already promised in writing
- Set the bar per class of product
- Name who gets asked, in advance
- Confident in either direction is the failure
basics
~20 sEverything about its own exposure: what it accepts into which products, what its existing contracts and licence obligations already require, and who it asks. The law is not the team's to settle, and stating it as settled is the error.
solid answer
~50 sSeparate the open question from the ones already answered. Whether model output attracts copyright, who would hold that right, and what counts as a permitted use of training material are genuinely open, are asked differently in different jurisdictions, and are not yours to decide — so say they are open. Most of the near-term risk is not there anyway. It sits in obligations the team already carries, which are settled and readable today: a customer contract warranting original work, a licence obligation attached to code already in the repository, a commitment about what ships in a product you license out. Those bind now, whatever happens to the copyright question. So decide the things you own: what you accept into which class of product, who you ask when it matters, and whether you could say which products were affected if the open question closed badly.
go deeper
Do not state this as settled in either direction, and do not repeat a rule you heard somewhere. Know that your team has a position, find out what it is, and ask before putting generated code somewhere unusual.
Explain why the copyright question is open while the licence and contract obligations are not, and why the second set is what binds the change you are making today.
Show that you separate what is open from what is already written down, that you know who to ask, and that you could say where you would look if the answer moved against you.
Own the bar for each class of product, the review date on that decision, and the records that would let the company answer which products are affected if the open question closes badly.
## Two questions that keep getting mixed together Ask most teams about the legal risk of generated code and you get one conversation about copyright, conducted with more confidence than the subject supports. There are really two questions in the room, and almost all of the answerable risk is in the second one. The first is **what the law says about model output**. It is genuinely open, it is not asked the same way everywhere, and it is moving. The second is **what your team has already promised in writing**. That one is not open at all. It is in contracts and licences you can read this afternoon, it binds today, and it is the risk most likely to arrive first. ## What sits on each side | the question | who can answer it | status | |---|---|---| | does model output attract copyright protection at all | not the team | open, and framed differently in different jurisdictions | | who would hold that right if it exists | not the team | open, and downstream of the first question | | was training on published source a permitted use | not the team | contested, and not uniform across markets | | does our customer contract warrant that delivered work is original | the team — read it | already decided, in writing | | what obligations attach to third-party code already in our repository | the team | already decided, by the licence on it | | what will we accept into a product we license out | the team, with counsel | decidable now | The useful discipline is to notice which side of that table a sentence belongs on before saying it. **Confidence is cheap on the top half and cheap for exactly the wrong reason**: nobody in the room can check you. ## What the team can decide for itself 1. **The bar for each class of product.** An internal tool, something you license to customers, and something carrying a warranty of originality do not have the same exposure, so they need not have the same bar. Saying that explicitly beats one rule that everybody quietly breaks in the low-stakes case. 2. **What you will not accept at all**, on grounds that do not depend on the open question. Code nobody can account for, in a component that warrants original work, is a problem under any resolution of the copyright question. 3. **Who you ask, and when.** Naming the escalation in advance is most of the value here. The answer is jurisdiction-specific and it is not the developer's to produce at the keyboard. 4. **Whether you could answer "which of our products would be affected?"** if the question resolved against you. That is a question about your own records, and it is answerable today whether or not anything ever resolves. 5. **When you will ask again.** A decision with no review date encodes today's uncertainty permanently, and this is a subject where that is a way of being wrong later without noticing. ## Why the confident answer is the one that fails The two confident positions fail identically. *Generated code is not protected, so there is nothing to think about* and *it is all derivative work, so we can never use it* are both settled answers to an unsettled question, and each replaces judgement with a slogan. An interviewer raising this subject is usually screening for that reflex, not testing your law. The answer that passes says what is open, says what binds anyway, and says who would be asked. ## Three things not to say - **Do not cite a case, a statute or a ruling from memory.** On this subject a half-remembered authority is worse than none, because it gets repeated by people who will not check it and it will be wrong somewhere. - **Do not offer one answer for every market.** A team shipping into more than one is asking more than one question, and the answers need not agree. - **Do not let "unsettled" become "nothing to do".** The obligations already in your contracts are settled, and they are the ones likely to bite first. ## The posture, stated plainly A team can hold a defensible position on a question the law has not answered. It looks like this: *we know which parts of this are open and we do not pretend otherwise; we know what our existing agreements require and we meet those; we have set a bar per class of product and written down who decides the hard cases; and if the open question closes against us, we can find out what it touches.* None of that requires knowing the answer, and all of it is lost by a team that decided the answer was obvious.
- What is the first thing you would check about a generated change that worries you?What the component it lands in already promises. Something you license out, or a deliverable under a warranty of originality, carries obligations written down and in force today; an internal tool usually carries fewer. That tells you whether the change needs a conversation at all, and it does not depend on the open question.
- Your team ships into several markets. How does that change the policy?It means you hold more than one answer, and an honest policy says so. Either set the bar by the strictest market you actually ship into, or segment by product and record which bar each one is on. The failure is a single confident rule written from one jurisdiction and applied everywhere.
- Does a supplier's assurance settle it?It changes who carries some of the cost if something goes wrong; it does not change what the law says or what your own contracts require. Read what it actually covers and what it requires of you in return, then carry on with the rest of the analysis. It is a commercial arrangement, not a ruling.
saying these in an interview costs you the question
- Generated code is public domain, so nothing applies
- All generated code is derivative work and must never be used
- Copyright is the team's main near-term legal risk here
- Counsel will give us the answer, so we need no position of our own
- One answer covers every market the team ships into
- The law is unsettled, so there is nothing to decide yet