Columbia India Hour Takeaways: AI and Intellectual Property - "Who Owns Innovation?"

As AI becomes central to how we work, questions of copyright and who gets credit for AI-generated or AI-assisted work are becoming impossible to ignore. The recent controversy over a mathematician's unpublished Navier-Stokes work being used to train an AI model only underscores the urgency.  In the third session of the Columbia India Hour series, Professor Shyam Balganesh (Sol Goldman Professor of Law, Columbia Law School) and Bahram Vakil (Co-founder & Senior Partner, AZB & Partners) took on AI and copyright explored the evolving legal landscape, the licensing deals now being brokered between publishers and AI companies, and what it all means for the future of the legal profession. 

September 28, 2026

Key takeaways from the session

AI companies building generative models are grappling with access to content,  much of which is high-value and copyrighted. The question is whether using that content, both to train models and to ground their outputs, constitutes copyright liability.

In the United States, as many as 147 lawsuits have been filed against AI model developers for copyright infringement. These lawsuits have not only questioned whether access to the content amounts to infringement but have questioned the very process of training the model, vectorisation of it. 

Different jurisdictions are encountering this in different ways:

What’s happening in India:

  • The Delhi High Court Benchmark (ANI vs. OpenAI): Asian News International (ANI), one of the major news agencies based in India, sued OpenAI, arguing its content had been scraped and used to train ChatGPT. The court ruled that training AI models on news content and making internal, private copies does not constitute prima facie copyright infringement warranting an injunction. 
  • One of the innovations of the court was that it took existing research exemption under Section 52A, which was primarily for human research, and dynamically interpreted it to cover AI model training. It treated the internal copy made during training as non-public and therefore not infringing.
  • Bahram pointed out that another question was whether the Indian courts the jurisdiction had since OpenAI servers were outside India. He said that the courts decided that it did fall under their jurisdiction since OpenAI has a huge customer base in India. 
  • The court leaned heavily into policy, emphasizing national interest, public purpose, and the push for India to become an AI innovation hub. Both speakers questioned whether individual courts are the right forum to make broad policy decisions or award economic subsidies to the tech industry— should courts be making that kind of industrial-policy call at all, or should it be Parliament's job?

What’s happening in the U.S.: 

  • Although many lawsuits are pending against AI developers in the U.S., only two major decisions have come out so far, both from the Northern District of California: Anthropic (Bartz v. Anthropic) and Meta (Kadrey v. Meta) cases. 
  • Interestingly, while both the decisions leaned towards “fair use”, unlike India, they didn’t talk about national interest, rather they cited the “transformative use" doctrine, which offers an exemption if one adds new value and purpose to a protected work.
  • However, the courts made a critical distinction between “access” and “training”. Anthropic had acquired copyrighted content through unauthorized, illegal "shadow libraries". This created severe legal liability. Anthropic faced a massive $1.5 billion settlement specifically over how it obtained access to the data, not how the model trained. 

Need for Legislative Reform and AI Regulation: 

  • Shyam noted that, much like the ANI vs. OpenAI decision in India, the US rulings were early-stage summary judgments. The courts explicitly flagged that more concrete issues may arise as records are better developed. Consequently, the law remains far from definitive, as even the legal community is still adapting to this new technology and its modalities. 
  • While current judicial sentiment leans toward treating model training as "fair use" due to the transformative nature of the technology, a central takeaway is that legislatures must modernize and amend laws for today’s facts rather than relying on the judiciary to stretch existing frameworks. Today’s statutory language was enacted long before AI existed. For instance, a recent Calcutta High Court decision (IndiaMART v. OpenAI) examined whether LLMs qualify for passive "intermediary" safe-harbor protections under Section 79 of the IT Act. The court indicated that LLMs act as "originators" of content because they introduce a creative or generative element rather than functioning as a pass-through—as Bahram put it, "they're not a post office". This classification raises the diligence standard AI companies must meet in India and underscores the urgent need for comprehensive AI regulation so that statutory amendments can guide the field without forcing courts to improvise. 

India’s Regulatory Dilemma

  • While India is eager not to miss the AI wave, policy discussions in Delhi aim to strike a balance avoiding both an overly strict framework (like the EU model) that could stifle innovation and a completely unchecked, "laissez-faire cowboy attitude". 
  • India is currently debating a comprehensive, horizontal regulatory model like the EU's EU AI Act (rather than the siloed approaches of China or the UK). This framework categorizes AI applications into three tiers:  i) Prohibited Areas: Strictly banned uses (e.g., child sexual abuse material). ii) High-Risk Areas: Regulated sectors requiring high oversight, such as healthcare and education. Iii) Low-Risk Areas: Broadly permitted "low-hanging fruit" designed to maximize innovation. 
  • Shyam expressed skepticism toward the EU model, arguing that a regulation-first approach creates heavy bureaucracy that struggles to adapt as technology rapidly evolves.  For example. platforms like Claude adding watermarks to comply with attribution requirements under the EU AI Act illustrate how bureaucracy forces immediate, top-down compliance measures. 
  • To boost its tech sector, Japan introduced a broad Text and Data Mining (TDM) copyright exemption years ago. Because no one clearly understood its legal boundaries, AI companies failed to invest, and the initiative stalled rather than driving growth. 
  • Shyam said that an “institutional symbiosis” was his preferred model. He emphasized that before drafting substantive rules, countries must determine which institutions they trust to govern. In the US, governance relies heavily on courts resolving issues case-by-case under broad statutory delegation. Shyam suggested India’s Parliament could set high-level frameworks while delegating case-by-case development to its vibrant judiciary. 
  • However, Bahram noted that relying heavily on Indian courts for regulatory development is impractical due to severe judicial backlogs, delays, and varying technical expertise. Bahram pointed to the UK’s principle-based model and regulatory sandboxes as a more workable middle ground, establishing core principles while giving businesses flexible environments to innovate. 
  • Bahram shared a recent example highlighting the murky reality of AI training data. A mathematician working on the complex Navier-Stokes equations accused OpenAI of accessing and training on his private research documents without authorization. While OpenAI countered that the AI simply solved the math problem quickly, the key concern remains how models obtain access to private, unpublished work. 

Licensing is quietly becoming the real answer

  • Both speakers agreed that the practical resolution to the AI-training copyright fight is emerging through the market, not the courts: a wave of licensing deals between publishers/rights-holders and AI companies.
  • According to Shyam's view, government's most useful role isn't to build a comprehensive licensing framework itself (he thinks that's unrealistic given how "hungry" frontier models are as they want to ingest everything, including "bad" content like Reddit posts, precisely because models need negative examples to learn to differentiate good from bad output)  but to lower transaction costs and create incentives for private licensing hubs/intermediaries to form.
  • He noted the volume of licensing deals is already large enough that it's "rendering the 147 lawsuits a little bit irrelevant" in practice, even though no clear industry-wide remuneration mechanism exists yet.
  • Bahram said that India is "a bit behind" on this but he already advises clients that licensing, not litigation risk-taking, is the practical way forward

Takeaways from the Q & A: 

Could a tool assess how likely a given AI product is to trigger copyright liability, the way risk models exist in finance/medicine/engineering?

Shyam answered that technically possible, but two big obstacles stand in the way today:no stable legal baseline yet and nobody outside the frontier labs fully knows how models like GPT, Claude, or Llama were actually trained or fine-tuned, and as companies shift toward closed-weight models, that opacity increasingly falls under trade-secret protection rather than copyright, making downstream risk assessment "so peripheral as to be almost useless." Something workable might exist in 6–8 months, but it will inherently trail the technology.

Can bureaucrats can be trusted to regulate something this technical? 

  • Bahram pointed to India's MeitY leadership and cited the UPI/payments success story, including Nandan Nilekani-style public-private partnerships, as proof India can build competent, technically literate institutions when it wants to.
  • Shyam added that, unlike the early internet era (when regulators genuinely debated whether software should even be classified as a "good"), today's regulators on both sides know they need technologically fluent staff. The real bottleneck isn't awareness, it's speed. 

AI, authorship, and Columbia Law's own policy

  • Shyam described Columbia Law School's internal task-force policy on AI use (which he helped shape): students may use AI, but the core value the school wants to protect is transparency/disclosure. This means not just disclosing that you used AI, but why and how, so the underlying thinking is visible.
  • On the "does AI replace lawyers" question, Bahram shared that young lawyers should draft their own points first then check it against AI output, then double-check the AI's output — never let AI think first.
  • Shyam referenced a 1990s American TV show, Doogie Howser, M.D. (about a teenage doctor), which he watched growing up in India, and quoted a law-school mentor's line: "there's no such thing as a Doogie Howser, JD" because lawyering depends on experience-built judgment and wisdom that, in his view, AI will never fully replicate, even as it takes over the mechanical/first-and-second-year tasks.

Authorship questions AI raises for IP law itself

  • AI cannot itself be recognized as an author or inventor. IP law's incentive rationale is built around motivating human creators, and a non-human agent doesn't need that incentive.
  • The genuinely hard cases, per Shyam, are heavily AI-assisted human works (e.g., someone spending 15 hours in dialogue with a chatbot to produce a novel or artwork). He's been advocating for treating this through a co-authorship framework rather than an all-or-nothing test. 
  • On whether fair use/fair dealing should just be replaced by a paid-permission or compensation requirement for AI training: both were candid that there's no clean answer yet, reiterating that the current handful of rulings are narrow and fact-specific, not general licenses for all AI training.

Closing advice to law students

  • Don't fear or avoid AI, build fluency in it as a productivity and research tool.  For instance, IP law becomes even more interesting layered with AI and India's incoming Digital Personal Data Protection (DPDP) rules, taking effect in June.
  • The single biggest mistake anyone can make right now is fear-driven avoidance of AI. Shyam explicitly rejected the "wait and see" or "stay technologically backward" stance, calling it a path to becoming "real dinosaurs," and framed the task ahead as learning to re-equip and work symbiotically with AI rather than resisting it.