Engelberg Center Live!

Open @ UNGA 81: Open Source Software

Episode Summary

This episode features audio from the Open S ource Software panel at Open @ UNGA 81: A Conversation on Collaboration. It was recorded on September 21, 2026.

Episode Notes

Open source has repeatedly reshaped technology markets by lowering barriers to entry, reducing rent extraction, and enabling innovation beyond what proprietary licensing allows. But open source code alone doesn’t guarantee competition, “sovereingnty,” or decentralization. A project can also be open while the surrounding market remains concentrated in cloud, compute, distribution, data, or complementary services, having little ultimate impact on the incentives of the actors in the market or the competitive dynamics. These tensions are especially visible in AI, where “open weights” are often treated as synonymous with open source even when other ingredients needed to study, reproduce, or operate the system remain unavailable. This panel examines open source as a market-shaping tool, asking what complementary conditions (architecture, protocols, procurement, governance) make it effective, and what outcomes are actually feasible in AI markets.

Episode Transcription

Announcer  0:01  
Welcome to Engelberg Center Live, a collection of audio from events held by the Engelberg Center on Innovation, Law, and Policy at NYU Law. This episode features audio from the open source software panel at Open at Unga 81, a conversation on collaboration. It was recorded on September 21, 2026.

Ilan Strauss  0:25  
So my name is Ilan Straus. I co-direct AI Disclosures Project with Tom O'Reilly. I'm an economist, which means this panel would probably be quite boring if it wasn't for the amazing panelists who are here with me. A big thanks to Creative Commons for basically organizing this panel and then allowing to call the AI Disclosures Project the co-host. So thank you for that. So the panelists, I encourage you to introduce yourself through the course of the discussion as you answer questions. Use examples from stuff you're working on. I think the questions we're going to tackle in this panel on open source AI, how is it shaping AI markets? Realistically, how can it shape AI markets? I think some of the questions we have already covered, but I think hopefully this panel we can bring new examples to the fore, and we can bring new theories. Hopefully, because all of your your backgrounds and the things you're working on are incredible, are unique, and I do think it's additive to what's already been discussed. So the goal today is to kind of think: well, how is open source actually shaping AI markets today? What are its potential, and what is what is its impact conditional on in the future? And hopefully we'll use concrete examples. Just I'm just going to give like two minutes just to set the scene. So on a definition, when an economist mentions a market, they often talk about a market structure. That's how you kind of define the market as an economist, and the structure are how many firms are there in the market? So is it a monopoly, like one firm, or is it very competitive with many? They care about that because that that has implications for the market's welfare and how value is distributed. And economists assume that the market structure is a function of the technology because the technology decides costs. You know, so if you have high fixed costs, okay, where you got like build a factory worth billions, chances are you know there's going to be fewer firms in the market because you know only so many companies have the capital need to sustain that fixed cost. Just like we saw with the old generation of tech, you maybe need a lot of servers, and then the other big thing on network effects, which was really relevant to the old generation of tech. Network effects also means that you can have a winner takes all or winner take you know a few firms dominate the market unless there are things which are sharing the network effect so everyone can benefit. But you know Facebook has is able to captive all the benefits. It's able to capture all the benefits of its users. Then only Facebook gets the the network effect and it gets the increase in average revenue. That's the market structure, but tech markets are kind of different, right? Because they're not so flat. There's a network. So what decides the market structure in networks? Well, we're still trying to figure that out. But in a network, we know things like standards matter to connect to the network, and standards matter for production because there's so many parts that that interconnect. You know, same when you make a car, there's lots of parts, but maybe tech markets are a bit different. So, with that definition in mind, I just want to say, give you two facts to consider in terms of AI market. Why it's so important to consider how to make this market less concentrated, more dynamically innovative in the future. So, in 2021, Apple's market capitalization-the value of Apple, its equity-was 94% of the German stock exchange's total capitalization. So it was almost one to one in 2021. Again, Germany is the powerhouse of Europe here. Even if like half of its employment and value add is not in the market, anyway, in 2026, Apple is now approaching twice the market capitalisation of the entire German stock exchange. Okay, it's at 1.7. So we see how there's growing divergence in in where economic power is concentrated within a handful of companies, and so in today's discussion, I'm hoping we can think about well, how can open source compete and change that balance of power? You know, one crude way I think about it, and hopefully you can raise this business models or business strategies.

Ilan Strauss  4:37  
One thing I think about maybe it doesn't fit Boeing versus Airbus. Boeing's an American company, and Europe at some point was like, well, we want to compete with Boeing, and it took them 30 years roughly to achieve parity with with the American company and overtake it, and they have different production models. You know how integrated they are, how they set the standards. You know Airbus, this pan Europe. Company, how they design the standards, how they facilitate R&D. So it's really about competing, not you know, not head on, but in a differentiated manner. You know, TikTok competed against Facebook in a differentiated manner. So I'm hoping we can raise some of those strategic questions. So when it comes to strategy, what should the strategy be of open source. You know, at the moment we hear a lot about open weight models. Okay, is this the right strategy? Because when you know I think about open source, I think about well, so many markets where there is open source. Android is open source. Chromium behind Google Chrome, it is open source, but the dominant company still can capture the value eventually, right? Android's open source, but Google still captures the value from Android because enforces a bunch of standards, okay? And then it also says you have to install our our Play Store where we, you know, we take cuts. So, is this really you know? So first, open weight models. Is this where we should be focusing on if we want to decentralize value to make it easier to enter the market? So yeah, maybe Amanda, if you want to kick it off for you know two or three minutes.

Amanda Casari  6:15  
Yeah, sure. Thank you. And so Amanda Caseri, I'm in the Google Open Source Programs Office. I do many things. AI is one of them right now. I think if the the question around the concept of even just centralizing the importance of models in a modern tech stack, I've been asking the same question really since 2022 when things started really gaining more attention in my personal life that is in my work life, which is what's new and what's different now. So I think that there was a con. The conversation was different in 2023 around models because the assumption was was that models would replace the entire technology stack. It would replace data extraction, consumption. It would move into feature extraction. Like the model was basically doing everything. The model itself was the product. It was not integrated into a product. It was not a feature. It was not an enhancer. It was not a flavor modifier. Like it was the actual product in and of itself. Three years later, I think we can find that that's actually not the case. I mean, it is a market in itself, but it is not the market in and of itself. So I'm really encouraged. By the way, thank you so much, Anna and Creative Commons, for structuring the conversation today around governance and infrastructure and data and hardware and so many other pieces of what it takes to build the modern infrastructure stack. And I want to double down on the question of like where we should be investing when we think about equalizing through software, modern tech stacks. I really appreciated Mitchell kicking off with the idea of interoperability and protocols and standards and global governance models, because I think in in a modern tech stack, this is an area, and I think Chris brought this up as well when we're talking around protocols of where innovation and and where technology can change in a way that we currently cannot predict or commodify right now, so I think if we're looking at one specific area of the stack and trying to see how we can bridge that gap or break that moat, what we're not identifying is where constraints happening and where are we going to allow people to speak a common language, both with technology as well as with human trust, that allows them to move in ways and come up with new ideas that we're not currently capturing through VC money. So right now, I think we are in an immature state with that. When it comes to the modern AI ecosystem, we can see that even the discussion that Jan brought up earlier, talking around really robots.txt and how that's a protocol. Protocol is not the same as a standard. There's no remedies associated when you break a protocol versus when you break a standard. Like you feel them through the remedies of how people are acting, but it's not the same as like having to go through remedies from a regulatory stance. I mean, we see that now too developing. I will say that there are some there are some protocols and standards, and I'll put quotes around those that have been donated to vendor-neutral foundations as a way of developing those to allow for new markets to exist. And I've looked at them and I've said these are markdown files. I don't understand how this is a standard. This is a this is a data structure. This is a metadata structure. This is not a standard. But the reality was was that there needed to be a vendor neutral place that different groups could come together and talk about how their products would interact and work with each other. That was not going to be waiting on standards bodies, quite frankly, because standards bodies just don't move at the same speed. So I think if if I was looking for where where we should be concentrating our attention, I actually do think there needs to be increased investment and work from the different groups into standards bodies, into protocols, into the vendor-neutral organizations like OSI that are actually working to try to make sure that access is not to the few but to the many, and at whatever level you're working on, that you're able to access that.

Ilan Strauss  10:00  
Very interesting. Well, hopefully we can drill down into some of these specific protocols or standards, like MCP, and then thinking about what's enabled, but what have the limitations been. So hopefully we can cover some of these specifics next. Madiha, hopefully I pronounce your name correctly.

Madiha Zahrah Choksi  10:17  
Yeah, good. I'm Madiha Zahrah Choksi. I'm an academic. I'm currently a postdoc at the Data Science Institute at Columbia. I don't have an answer for MCP, but I do have something to to maybe add to Amanda's points. And I think that's fundamentally. I don't think that power moves around randomly, right? I think it moves on purpose. So this idea of like commoditize your complement, right? Like Google didn't open up Android and Chromium for fun. Like it was in service of search and ads, right? The Llama family of models were released as open source as a competitive advantage to OpenAI and Anthropic, right? We see this happening over and over again, and so this kind of idea of opening a layer in the stack will fix everything. I don't necessarily think so. I think this is a move that's more on a complement, right, in service of some kind of competition. So, if this this question around where should we concentrate or focus our efforts, you know, I'm an academic. I study communities and and user groups, and so I think that's where the layers connect, right? Like, where do these layers connect? Where does participation actually happen? So for AI, that's model hubs, right? That's model scope. That's Hugging Face, and I think that one of the most kind of interesting but also terrifying research studies that I've taken up over the last few years has been on Hugging Face and and what does community participation in open source model ecosystem actually look like and what it actually looks like is this right? It doesn't look like this, which is what we previously observed in open source software ecosystems. There's like a significant power law distribution in open source model ecosystems, collaboration and community efforts. However, model scope looks a little bit different, and the innovation coming out of China's communities looks very different, so maybe I'll I'll stop a little bit there. But I think that we need to scrutinize and think more critically about the model hub spaces, right? And I think I'll just add one more thing, which is on Hugging Face there is a page that lists licenses that a community can add to their repository, right? It's like license like licenses for repositories. It's just a static web page of 86 licenses. You cannot click into them. The Creative Commons licenses are there, yay! But you can't click them. There's no definitions. There's no documentation. So what does that actually mean, right? GitHub, it looks a little bit different, but on Hugging Face, this is this is currently as of two years ago and as of last night, right? True, you can you can see this.

Ilan Strauss  13:08  
Interesting. And so, what is a you said model hubs? What what is a model hub?

Madiha Zahrah Choksi  13:12  
I model hub is just like a term for like where open source models are kind of building, like like the GitHub or so the Hugging Face or Model Scope for open source AI development, model development.

Ilan Strauss  13:30  
Okay, good. No, that's good to know. It's like a term.

Ilan Strauss  13:33  
Yeah. Term right. No, good. Just everyone on the same page. Yeah. Any other panelists want to add anything on that or

Ben Moskowitz  13:42  
what you? you asked about MCP, so I just say a brief word about MCP. Sure. So MCP is the model context protocol. It is de facto the way that applications built on AI connect to other sources of data to the applications. It's fantastic if you're a developer, but just speaking to the limitations of an open standard or an open protocol to create competition. It's called the model context protocol because built in is an assumption that we need to make it easier to bring the data to the models. In fact, there's nothing like a marketplace like you have even with the App Store, which people critiqued for being you know a closed garden and so on. MCP has not created a marketplace or anything resembling a marketplace on ChatGPT or Cloud. If I want to bring a new application to market on top of those things, treating them almost as a platform, I can't. And so we can get more into the market structuration and what's possible. But I think MCP is a great example of a very neat protocol that's not doing anything to really create new competition dynamic,

Ilan Strauss  14:44  
but okay. So MCP, so it's like you know with Claude Code because of MCP, I can now get emails from my Gmail, right? Potentially with read or write access, you know. So it's like this universal adapter for APIs, but. But so, doesn't it mean anyone can create an application then, you know, potentially, and then you can integrate your AI into that application in the enterprise

Ben Moskowitz  15:11  
market? Yeah, again, the enterprise market for sure. But in the consumer market, MCP doesn't do anything to challenge Cloud or to challenge ChatGPT because my account is with Cloud or ChatGPT, Cloud ChatGPT would love to have access to your email and your documents and so on, but they're not doing so in a way that creates the possibility of like new consumer entrants to build value on top of those platforms yet, and that might change. I

Ilan Strauss  15:33  
see. So you're basically saying you're still reliant on the same platforms to call those applications, right? Is everything? But does it? I mean, risk thinking aloud here. Does it not make it easier, though, then, for you to make your own clients? Because then, you know, instead of me, like, you know, if you think about an effective AI agent today, right? Imagine there wasn't a standard like MCP. I might need a bespoke deal to make my agent relevant. You know, I would need a deal with Google, you know, maybe, or I would need then to hardcode an integration to Google's API and then to Slack's API. So it isn't easier in some ways to then make these harnesses to spin them up because you know you don't, you know, you're not starting from scratch when it comes to integrating your harness with these tools. So doesn't that reduce barriers to entry at that layer at least, or not?

Amanda Casari  16:24  
Please. So I think one thing that we we should make sure we're adding in here is when you're talking about ease and and who that ease is for. In the past few years, one of the concepts that has fundamentally changed when we're talking about an open internet is is having an interaction, and what are we assuming that that entity is? And so, in because where this protocol sits as part of a stack is conflates the idea of a probabilistic versus deterministic process. And so, part of the problem we're having now when we talk about trust, like little t trust and and humans and machines and the ecosystem we're creating, or we can't conflate that with deterministic trust, which modern security is built on. And so, by doing that, by having machines act as users in order to gain access to these systems, to be able to gain access to APIs that we do not have processes right now that where we can understand coming forward into a machine, is this a bot? Is this an agent? Is this a machine? Is this a human? We're catching up there, and so the ease and for whom and whether or not your enterprise or whether or not you're consumer right now, a lot of those have broken down into a way that is combining and conflating those assumptions, and we haven't fixed that yet, and it means take away from the point you're trying to make about MCP. But I feel like we have not talked about that here enough. That there's like the concept of trust on the internet. It looks fundamentally different now than it even did a year and a half ago.

Ilan Strauss  17:55  
And so, are you saying that in effect this is stopping markets from arising, and markets could be, you know, like privately run, like an app store, but still at least allows potentially for broader participation and value, right? Same like if you monetize your content on YouTube, or you're also saying that these open standards, right, in comparison to these private markets, they're also being impeded from arising to create, you know, these distributed layers of the market because there is a lack of trust, security, identification, verification. Sir, has

Amanda Casari  18:32  
it took over? This is such an

Ben Moskowitz  18:33  
interesting conversation because, and just building on something that Amanda had said, maybe at the dawn of LLMs, we thought, oh my God, these are going to be the Unisy systems, right? Everything's going to run through an AI model. And now, a couple years later, well, that still may be in the long run, but we can see that it's a much more complicated situation. So, you know, there's the question of is the power in the market from the model, or is it from the point of distribution where a customer or a user interacts with the model. If I'm in the let's just say the consumer market, a lot of the value comes from the harness or the client that connects to the model. So EM made a great point, or at least a wish list item in the last panel, saying, "Wow, the open world, the open source world, should have a go-to-market strategy for an open alternative to buying a ChatGPT subscription and so on. Well, the models that power that kind of thing are becoming commoditized. In fact, you have many choices for a model that is 80, 90, maybe even 99% good enough for the things that most people want. Whether that's an OpenAI model, an Anthropic model, a Google model, the long tail of open models from companies you've never heard of, quants, distillations, open source models that run on your computer. The model is no longer really the unique thing for most people. It is driving towards a commodity, and yet in any market, the ability to distribute. Distribute the application is very important, and you know if we're thinking about the consumer market, even though technically we could set up very easily, and there are open alternatives, they're not yet distributable, right? I mean, so the distribution advantage, the advertising advantages, and so on, the data and the agreements that they have access to are different, right? And so it begins to bleed into all the other discussions that we've had today. So I applaud you for setting up a really broad market structure conversation because we're at a moment where, one way or the other, all of the markets, consumer enterprise, are being questioned. Interesting. Again, that I guess you're talking

Ilan Strauss  20:36  
about there about distribution store matters, especially for the consumer. You know, like my mom or other people. You know, you're going to go on Google. You type into Google search, and then there's an AI and an embedded. Same with Meta, right? Their embedded AI usage has gone up. They're saying distribution matters, and that seems like a really important point. Dwayne, you want to say where you think the focus should be now that maybe we've you know we've solved models, right? So we can move on to something else.

Speaker 1  21:01  
Yeah, yeah, we fixed it, right? Yeah. Actually, the thing I want to talk about is a is a theme I hear coming up in in some of the questions and responses. Is open going to do this? Is open preventing that? And so on. Open by itself does very little. It enables a lot, right? And so if we're looking to open source just as a methodology to create markets to push protocols that everyone can contribute to, that's not enough. And I think I believe that's sort of the point EM was driving to that there's no particular go-to-market strategy, and that's not really what open is about, right? You can't have a lot of these things without open. You can't have transparency. You can have reproducibility and so on. But we, especially those of us who are close to the space and love open and sort of embrace it as our as our raison d'être, like we we we lose some lose sight sometimes of the fact that it's not enough by itself.

Ilan Strauss  22:00  
So, what do you think that other than you know? So, if it enables, so what you know in Europe, let's say, does that mean then that Europe and America we should see a flourishing of innovation across the stack because we have open models? Therefore, it should be enabling distributed innovation down the stack, or is things still missing?

Speaker 1  22:23  
I think you can see them, but yes, I think there are things that are still missing. I think Airbus was a really good example to cite for a lot of reasons. It didn't happen organically; it happened as a result of significant, long-term, sustained investment. And I'm sure 30 years ago, it was inconceivable that anybody could get to the level that they could compete with Boeing because of their hold on the market. Just as it was inconceivable at one point that anybody could disrupt Unix, it was inconceivable at one point that anybody could unseat Internet Explorer. Right? All of these things we forget, especially, and it's so hard right now because the space is changing so hard. The big players are so big that that it can feel sometimes like it's it's just not attainable. There's a there's an adage that the the best time to to plant a tree is 30 years ago, 20 years ago, 10 years ago, right? So if we if we want these kinds of markets to flourish in in in Europe in anywhere in the world, you have to invest in creating those markets.

Ilan Strauss  23:30  
100% I wouldn't mind taking five minutes just to see if we can drill deeper into the relationship between open protocols as an open standard for communication on a network and open source software, because you know, with an open standard, right? I mean, as long as you are talking that language, you can be talking it as a private piece of software or a public open piece of software, right? It's agnostic, you know. So, like, I can connect to the internet now, you know, with kind of any kind of hardware, you know, most kinds of software, as long as I'm talking those specific internet protocols, and so it kind of allows for continued innovation potentially on the back end, as long as the standard isn't set in too much of a rigid way, but I thought we'll see which of the panelists want to tackle this. We can go from left to right, and to be like for the AI stack, you know, where do you think the most important interventions will come? The balance between open source and open protocols. Do you need both? And maybe you know, just if there's a concrete example that you want to discuss, whether it's a protocol in AI or a model in AI, you know, that'd be great. So we'll go from left to right. You don't have to answer this one. There'll be other questions.

Madiha Zahrah Choksi  24:53  
Okay, come back to me.

Speaker 2  24:56  
Cool.

Ben Moskowitz  24:56  
Well, I think it's constructed to talk about a stack. And you know, I at the same time, if you'll permit me, I think we should also kind of step back a little bit, because there's discussions of what protocols are needed for the agentic internet, for instance. Right, there's a whole set of them. There's discussions for yeah,

Ilan Strauss  25:15  
just like the trust or identification that could be a protocol, right? Yeah,

Ben Moskowitz  25:18  
you've got not just MCP, you've got agent-to-agent protocol. You have agent payment protocols. You have a flourishing of protocols, and they're very special purpose. And in the enterprise space, they're going to get worked out. And you know, I feel like that's it's a very healthy space, right? At the same time, if you if you think about the market for AI being well, how do consumers, citizens, people access the power of AI? You need a slightly different framework, and so you know this is more like how do we build a stack where you can enable lots of the knowledge needed to deliver a really advanced system like a ChatGPT? How is enough of that knowledge public and actionable, so that the course of innovation will will take its course, whatever that may be. And so the stack is a is a very helpful organizing framework because you want enough knowledge in the commons to exist on how do you train a model. You want exemplar models. You do want open weight models that you can pick up and play with. You want harnesses and applications that you can build on, you do want protocols. You need a whole stack to enable the kind of openness I think kind of unites a group like this. But also, it's interesting because the internet is a very constraining way to think about what's coming. And I, you know, it's said that if you're born after January 1st, 1970, you've lived your whole life in Unix time, okay. 50 years plus, we've all had you know these mental models for the governance of technology and how rewire society and all that. But the people building AI, you know, in the early days of the LLM revolution, were saying this is bigger than the internet. This is like electricity or fire, or the written word, right? And so I think when we talk about openness, there's like a degree of there's a value alignment here that's bigger than how do we keep it open, like the like the internet. Yes, I'd like AI to be more like the web or email than a railroad, privately owned and operated, vertically integrated. You know, yes, I'd like the AI writ large, be more like the email and the internet, and not an oil company, and that creates a lot of inequality and amplifies inequality. But you know, I think we need to think more broadly about how openness across the stack enables choice, competition, and even freedom, if you'll permit me. And so you know, protocols fantastic, but but I almost think that the stack is the right way to think about this. The knowledge, the techné, you know, for you to have the ability to create something as amazing as ChatGPT. What's it going to be in 510, 1520, years?

Ilan Strauss  27:55  
100%

Speaker 1  27:56  
I I know you prefaced that question by hoping that we would get specific, and I'm going to disappoint you. I'm sorry, but I'm also going to answer the question through a slightly different mindset. We're talking about, you know, where can we pick the right place in in the stack or in the protocols in order for this to work? I don't think it's about where. I think it's about who, right? Who can you mobilize for collective action for collaboration the most successfully, and it's not about finding a leverage point that's the most likely thing to be open. It's about creating that leverage point through that collaboration and collective action. So if we if we all agree that you know there's a particular if if the harness is the place, right? If we agree that the harness is the place where the most leverage is, and nobody comes to work on it, it doesn't matter. We have to find the place that everyone will will collaborate in order to create those leverage points.

Ilan Strauss  28:53  
And that seems to be a very open source forward answer, as far as I can tell, because in some ways, open source is a different kind of potentially institutional framework, you know, in terms of how people work across borders, how they collaborate. Yes, open source-it's not free, as was noted earlier. Even if it can be free to to distribute, it can require lots of money. Whether that money is from, you know, third-party firms who are private firms wanting to disrupt the incumbent, like that was the case with Anthropic, who wanted to, you know, develop this MTP standard to to disrupt some of OpenAI's the incumbent their standards. So, I mean, can open source as a organizing not just principle but kind of like structural form. You know, I mean, is that then the the right way to think about it in terms of how people are going to be mobilized, or how, you know, I don't know, these middle countries can can band together. But again, you know, open source is is global and. Cross border, so it's still it's hard for me to kind of square this cross border internationalism with states and these kind of borders, you know. Whereas like bits and data is, you know, it's kind of going everywhere.

Ben Moskowitz  30:14  
Well, so I'll pick up on your question of is open source still like a helpful construct? Well, of course it is, right? Of course it is, but also it's under some pressure, and I think OSI can speak to this. There's a definition of open source AI, which is very robust, and a lot of what we call open source AI isn't strictly speaking open source. Now I'm okay with that because I think that open weight AI-that is, you release the model people can use it. In a lot of cases, you get a ton of the value of open source, even if it's not open source all the way down to the studs, right? And I think that you know, at least when I come to these discussions and I think about how we're going to have an open market, how are we going to have open innovation? I think that we do need to think a bit more broadly about open and not just open source. I'll give you another example: distillation is a hot topic. The distillation is when you train another model on a really powerful model. Now, if you're anthropic, you complain that all of these Chinese groups are building these models by stealing your model outputs by doing distillation on you know Fable and things like that. And I think that's awfully rich if you trained your model on everything on the internet, right, and then you go off and say, "Oh, but you can't train your model on my model, as an as an open person, my construction of open is cuts both ways, buddy, right? And I think that this is not an open source discussion, but in spirit, to me, it is. You know, I feel that we need to have the knowledge, the techné. I think we need the ability for people to study and inspect and create new things with knowledge and data and systems to be preserved in the AI era. To not recreate these black boxes that are highly governed and you know AI safety is a big topic, but if AI safety means that we need to turn away from a thesaurus, I think that's actually you know not a good thing. That's something that we want, and so I almost go back to something that predates open source. And there's people here who will laugh at me for bringing this up, but you know, before open source software was free software, right? And open source is sort of v2 of free software. We may be entering into an era where we need a v3 of free software that encompasses things like well, what's the right dogma about distillation? Right, reasonable people can disagree.

Ilan Strauss  32:27  
No, I like this. It seems relevant to consider you know more Cory Doctorow style like you know the right to be adversarial here and you know backwards engineer things maybe. But I saw Madea. You might have wanted to say something in there, and then we can hear from Amanda on protocols versus open source.

Madiha Zahrah Choksi  32:46  
So, a couple points I just wanted to to agree on, kind of complement that you were highlighting, Ben. So, the first is the open source AI definition, and I think that one of the challenges that academics have been having with the definition and with kind of researching in the space or interviewing or kind of looking into what open models or how startups are using open source models or like what the licensing structures look like. What's complicated here, or what's okay for us to even use, and what's complicated here is that the definition is the definition of open source AI. I think is is on its way, but I think that there is one part that that is a little bit difficult to follow, right? Like we are we are trying to apply the ideology of open source software and free software to a technology that is fundamentally materially different than what open source software looked like. Right from its development to its implementation, these are different artifacts that we are dealing with. So when we have this open source AI definition that is a little bit, I think I would argue focused on availability of training data code, reproducibility, but also on its downstream use, we are kind of forcing together two parts of this definition into one. Right? How do we reproduce the model bit for bit, and how do we allow for it to be reused? Those are two different, I think ideologies that we need to kind of find a way for them to to work together. And the point about distillation, I think, is a is a great one. And I think just this morning there was a really great blog post by Nathan Lambert that talks about distillation as well. And what he says is that distillation is not a great explanation for all of the success that the models for the China the family of models coming from China are are being kind of attached to the story. It's actually maybe at best explains like a performance gap, but that's about it, right? And again, on on. Kind of research and academia side, the Quen family models are like a foundational layer of research. So, in machine learning communities, communities that I'm a part of, you are benchmarking on all of these models. Like you will not get accept your work will not be accepted to a conference unless you show your work against these models and these benchmarks because they are that relevant.

Ilan Strauss  35:22  
Interesting. Yeah, I think those are some important points that help also differentiate parts of the stack. You know, once you get inside the tech, Amanda, I'm curious to hear from you, and maybe in your answer, I'd love to hear what you're working on at Google in terms of open source. I'm sure the audience is also curious.

Amanda Casari  35:39  
Sure. Yeah. I think first to address the question of fundamental change and that question around kind of like open weight models and whether or not that is what's necessary to distribute value and market power, whether it's protocols, whether it's some other aspect. I think one of the things that we we haven't talked about enough, and I think Madia actually first started addressing this was the concept of labor and skills, and so right now there is a super high concentration and assumption about the necessary labor and the technical skills and the the kinds of education that is necessary to build technology stacks like this, and I would challenge whether or not that's true. I think, as somebody again who you might be surprised to know that I am not always a hit at parties when I start discussing technology, especially when it comes down to the question of when people say things to me like pre-training and post-training and fine tuning and I look at them very clearly and I say we're talking about feature engineering right like this is the aspect of information science we're discussing and the fact that this is the quite frankly like most hated class of most computer science education but what they're really talking about is taking data putting it to a way that statistical processes can be able to linearize and pass through, and they don't like that because it turns the value of their labor into something that is worth much less on the current market. Quite frankly, you earn much less money as a data engineer in traditional ML processes and companies than you do as an AI engineer or somebody who focuses on pre-training and post-training. So companies who have spent a lot of money on their labor and on their intellectual property and on their knowledge in that area have a reason to make sure that that investment is protected. And part of that is by making sure that that's highly valued as part of your IP stack, as opposed to discussing the question of like, well, why, while all this stuff is going on, you know, maybe a group of open source developers or people or engineers or folks who are familiar with technology stacks can take a pile of data, use some open source open access data, use some open source software. They can put it through commodity hardware. Turns out you can shrink it down. You can link some laptops together, things that are already existing, and make something that is high quality to solve good enough problem spaces on good enough hardware. Does that actually solve the problem you need? And what's that worth? So I think then it's the cost of the problem solving comes much down, much farther down, which is probably frightening for some people who have invested a lot into that space. But I think that question of like, are we at the point where, as a society, we're asking how much should be invested here, and how much are we getting back from that, and what's the right size for a use case on good enough hardware. I think we're approaching that more rapidly than I think is is showing up in the headlines.

Ben Moskowitz  38:46  
Yeah,

Ilan Strauss  38:47  
it seems like that you're touching on maybe the architecture of the system, right? So open or closed. It's you know if you can run this thing in your laptop, that's kind of a different architecture, I guess, maybe to one big centralized model. We were going to discuss two other topics, which I thought was interesting. One was like the cloud. Basically, you know, over time, if we expect inference costs to grow, even if we're still using laptops and mobile, then that'll be a source of kind of profits or rent extraction in society that grows over time. And the other was, well, if you may waved your magic wand of industrial policy, what would you consider? But instead, let's go to the audience for questions and comments, and then you guys can always close with someone like that. I see Dwayne wants to say something. Go for it. Yeah.

Speaker 1  39:30  
While the audience is like queuing up their questions, because the open source definition came up twice in conversation, I didn't want to jump in. And while everyone else was responding, I also forgot to introduce myself when I started talking. I'm Dwayne O'Brien. I'm the executive director at the Open Source Initiative. We maintain the open source definition that describes if a software license is considered open source or not. Four years ago, we embarked on a process to craft an open source AI definition that was published two years ago as the version 1.0. It's not without criticism, and a lot has changed since then. It is. It took something like 10 years for the open source definition to stabilize into the set of principles it is now, and this is a continuous and ongoing process. So, if you want to be involved as a participant in this next round of discussions, yes. Please let's talk, and we're going to begin those conversations very soon.

Ilan Strauss  40:26  
Thanks for that invitation to the audience. Okay, so any if you have any questions, or if you just want to make a comment on on on some of the things we've been discussing, now is the chance, and I'll give you the mic. Hi,

Speaker 2  40:46  
I'm Jim Furkerman of Tech Matters Technology for the nonprofit sector. A lot of discussion of open weight models. A lot of people I talk to in the sector are kind of terrified of Chinese open source models and other models as having no idea what's in there, whether it's security risks or vulnerable people having their data being exfiltrated. How would you know? So is that something? Why does why do very few people seem to be worried about this? Or maybe they are, and I just don't hear about them so much. But I get the impression like almost all of industry is like now 40% open weight model. So, what's going on here? Why shouldn't I be worried? Good question. Should we just should we just quickly see if we can get one more

Ilan Strauss  41:30  
question? We'll take them all together. Is that okay?

Speaker 1  41:31  
Yeah, that's right.

Ilan Strauss  41:32  
Let's just go to the back there. Okay.

Speaker 3  41:39  
Kaurav here from Civic Data Lab India, you would have heard this. Like a lot of countries are now gunning for model sovereignty, AI sovereignty. How do you see open source and AI sovereignty coming together? Because open source fundamentally means level playing field, more collaboration. How do we balance that out?

Ilan Strauss  42:01  
Thank you. Okay, so then let's just take those two questions. Yeah, who wants to start?

Speaker 1  42:16  
I'll start with the open open wait question, and and would instead ask why are you not worried about the things you can inspect? Not at

Madiha Zahrah Choksi  42:25  
all.

Speaker 1  42:27  
This this entire process involves offloading, offloading or owning a certain amount of risk. Right. If you say I don't have to worry about this because Anthropic wrote it, you offloaded all of the risk to them, and if you do that without ignoring your own risk posture, where you deploy things, how you deploy things, how you vet things, how you make sure that they're doing what they're supposed to do, you're also making an error on your part, right? So, open weight models they give you some ability to to tune and tweak and and so on and so forth, but the more open the technologies are, and I would I would say that it does matter what we call open source and what we don't. Sort of our whole thing. The the more open these things are, the more you can inspect them. But if you are unsure about something, don't give it access to your email and credit cards, right? Start by start with the process of constraining risk from your side.

Madiha Zahrah Choksi  43:25  
I think I would just add to that. I think this morning there was a panel, and one of the speakers said, you know, we have to be really careful about the narratives we hear, and think about who's saying them, and like what incentives they have for us to kind of harp on this narrative. Like, oh, AI, super AI should we be scared are we going to be around right like all of I think that a lot of the rhetoric around the models coming out of China are unfounded at least in research communities and academic circles we are all running the the open models from China like we are researching with them we are testing with them. I think that a lot of academics and researchers would argue that the model families, particularly Quen, are the leaders in open weights, like as it currently stands, right? And that this is there may be two months behind the frontier models, like out of the U.S. so there are these I think narratives that American Frontier Labs would like to push towards us to kind of win the so-called AI race. But I think that there is a lot of really interesting rhetoric out there, and particularly if you were to look into the model, like the small model spaces, small model communities. So our local llama is like a great place to go to read about what people are doing with hardware, extra compute, GPUs, like a like a gaming cafe, right? Like people are sitting in communities and fine tuning. Models and and kind of discussing constraints or against both sides, Chinese open weights models and American. So, I would say like the the kind of small model spaces also one that might give you some answers and kind of bust some of these myths around.

Ben Moskowitz  45:17  
Yeah, and maybe just to take that opportunity to pick up from what Dwayne and Media both said, and and speak to the sovereignty question. So I think if you're a nonprofit CTO, you can use Chinese appointment models. It's fine. It's so fun. If you're worried about security, you know, go to any one of the companies that will handle the inference for you. Check out their security. They'll serve you GLM or Quan or any of those things, no problem. Now, if you're worried about sovereignty, right? If you're working for the government, you know maybe you have some different strategic objectives, right? Well, then an open wage model is not going to be a sufficient investment or strategy, right? You may want a sovereign model, and that might be part of your export policy. That might be part of your regional influence policy. You know, Indonesia has this incredible sea lion project that's trying to equip all Southeast Asian dialects with a foundation model they can own. There's groups like the Allen Institute where you can inspect things, you know, down to the stems again. You know, and and if you are you know thinking about U.S.-China competition and you're a middle power, you know, country, and you don't want your future to belong either to access to a frontier model from China or the U.S. Then yes, you you may want to think about sovereignty in a way that says that open weights is not sufficient, right? But I think if we can distinguish between the needs of like a smaller organization looking at cost, well, I have some generalized tasks, and these tokens are 100 times less expensive. I think that you can feel pretty good about open weights most of the time. Now, that's not the right answer if your consideration is national sovereignty, right? And I think they're something like, you know, is your destiny depending on some actor that may not share your values? Then it's a different discussion.

Amanda Casari  47:01  
If I was to bring these two together as well, I think that fundamentally this is a question of who on the internet can I trust if that's not me. And one, and I feel like asking about some of my work at Google. One of the remedies or solutions to this that I have been pushing back on, and we collectively have been taking a stance against, is the concept that you must have real identity on the internet in order to contribute to open source. So that's something we have taken a hard stance on, especially as open source security has before the AI initiatives were pushing. There, if anybody else remembers the open source security push that came through in 2020 1920 23, and but the concept that of who you are on the internet does matter. Who you are on the internet that's contributing to open source, and there has and is still a question for attestations and bill of materials, and around sovereignty, and where high-risk systems get very concerned of who has access to that and who's contributing to that. But then I think what we as a community, and we as scholars, and we in industry have to assess is that who gets to say who somebody's identity is, and who holds that power, who gets to give any identification documents, and who who gets to hold that trust of who an identity is, and part of the openness that makes the internet work is that trust is not built off of a concept of true identity that's held by any one particular governmental organization. So we have to build trust mechanisms in other ways, and we have worked for decades now, to be able to do that, and so I think that when it comes down to how do I, who can I trust, how do I build systems that I can trust, there's ways that we work together to do that. But trying to make sure that we're making things most secure by locking things down on individuals and putting trust in certain vectors, we should continue to challenge that-that whether or not that's going to solve the problems we're trying to, or create more problems that's going to reduce access, reduce approachability, reduce the ability of the global south to participate, reduce who is allowed to be on the internet and the access to the information that they can participate in.

Speaker 1  49:16  
I, I love that you brought this up for so many reasons, this this conversation about hard identity on the internet-it's not about trust. It's about who am I mad at, right? Who do I get to be mad at? Who am I going to chase down and try to punish? Who do

Amanda Casari  49:33  
I get to control?

Speaker 1  49:34  
Right. And ultimately, if if any development process requires me to personally know who you are to invest trust in you as a person. There are fundamental problems in the rest of the process. Right, you have to be able to trust that if code comes in, no matter where it came from, that if it has issues, you have enough eyes on it and enough processes put in place that you can find those issues. That. That it's inspected well, that everything else is working, and we see a reaction as a as a result of AI generated contributions across some open source communities, where they've closed the contributions or they're going back to just the people that they know, which is an understandable reaction. But it also highlights the same problem if we've entrusted trust all the way, and just like who a person is, we've we've missed the mark.

Ben Moskowitz  50:24  
Can I add to that too? Just because this is the open source discussion, wow, AI is really challenging the economics of open source in some different ways. So you know, there's the question of I have to deal with all these slop contributions showing up. Forget about bots, even right. I've got a bunch of mediocre people who know how to vibe code a pull request now, and I have to review like 100 pull requests, and only like five of them are actually any good. So that's where AI is changing the economics of open source in ways that put more pressure on maintainers. Another one that I think is really interesting is the refactoring discussion. So in the past, you know, if you had a disagreement about the the maintainers' kind of vision for the future, you might fork right, and then if you bring enough people along with you, now you kind of control the momentum of the project. But now I can kind of point at the at the project. I can say, help me re-implement this thing, soup to nuts. And in some cases, now there's these these questions of can I then apply a new license to a thing that's inspired by an open source project, right? A license that I might not have chosen for the project. So the cumulative, you know, decade, let's say, of work that's gone into making this open source project useful, now somebody can re-implement it by refactoring it, call it a new project, and give it a new license. So I think we haven't even begun to scratch the ways that AI is going to change the economics of open source. There's some deeper questions.

Ilan Strauss  51:41  
No, for sure, that's a big one. And by the way, when do we have until one minute? Oh, I see. Okay. Well, then we can consider those the the closing comments. I don't know. Okay. So then everyone can have a closing comment, and then mine will just be responding to the interesting question about sovereignty, and just to say again, David Ricardo's chapter on machinery. He was a board Jewish accountant, and he said, in international markets, you can be sovereign, but you can have no autonomy because you are a part of international markets, and so you have to be innovative. And then sovereignty and autonomy is much easier when the technology is more ossified and commodified, right? Like you can deliver water, okay, and you can have some kind of, and it's fixed. So you know, in in the AI markets where things are commodified, there's more room for sovereignty. I would argue, you know, so that's where the cloud and stuff like that comes off together, but otherwise, most sovereign projects fail because they they struggle to remain innovative. You know, so I guess it's about if you're focusing on a commodified or a non-commodified part of the stack. So let's do closing comments. Yeah,

Madiha Zahrah Choksi  52:56  
sure. Closing. I just want to add to something that Ben and Dwayne were saying about the maintainer burnout problem. I mean, this is like a huge, huge problem. I think that one of the one of the ways that I would tie it together to what I started with is governance, right? Like one of the reasons that maintainers are struggling so much on GitHub vibe-coded AI slop contributions is because GitHub has not provided tools for maintainers to adequately manage all of the AI contributions. Right, there are these norms that are emerging on GitHub. For example, people are adding the little lobster to say like ah this was vib coded right without saying it in so many words or you have contribution guidelines or contributions.md or agent md files emerging whereas in the past you would see license files or READMEs that would have that kind of policy around contributions articulated but these are now kind of emerging in like these like very disorganized ways because they they don't know what else to do, and so I think I started by talking about hugging face and the kind of mess of licenses. But all of this kind of really does come down to what does contribution look like? What does participation and collaboration look like? And and like in service of innovation, but that really does mean that things are breaking in this really scary way. Yeah.

Ben Moskowitz  54:28  
Yes, agreed. Reasons to be helpful. There's never been more open knowledge. There's never been more open collaboration. There's never been more open innovation. So things are good. At the same time, there's never been the same potential for enclosure. There's never been the same potential for powerful gatekeepers, and there's huge potential for inequality across all these things. And so we're at a moment of incredible opportunity and challenge. And groups like Creative Commons are incredibly crucial to steering us towards an open society. Thank you. More hopeful.

Speaker 1  54:58  
I don't know that I would say never. Right. If we go back to Victorian times and look at capture that was happening in the first industrial age, like it was significantly different than what we are now. And the reason I'm I'm not picking on you, but the reason I'm picking on that point is because if we if we internalize the narrative that this is our only chance to do something, we undercut our ability to think positively about it in the future, right? I've had many conversations with people before I took the role at the at the OSI who were very stressed out about the process of the open source AI definition because of the moment in time that we were in, and those moments keep coming back, right? Yes, there's a lot happening. Yes, it's changing faster than than any of us can really keep up with. Well, most of us can keep up with agents, and yes, the big players feel enormous, right? But we still have the ability to enact these changes, and they just they keep coming back over time. So let's not say never, because it's important, but don't sell ourselves short for the future. I guess.

Amanda Casari  56:08  
Yeah, I think I think one thing I would hope that this group especially can keep in mind is that there's this terrible communication pattern a lot of tech has gotten into, where they just throw like like what they call now a spec, which used to have specs. This is not a spec, but they throw something at you, and they're like, "What do you think about this? and and it's very overwhelming, especially when you get too many of these at once. And so, actually, a lot of my work now is to take a breath and to sit down and talk with engineers and teams and say, so what problem are you trying to solve? And to really move from the question of here's the here's this thing I made, what do we do with it? Into is this problem something that is valuable, well scoped, worth solving, and is this something that can be solved with information system techniques? So I would say that right now, again, like talking about where it's possible to get distracted right now, I think it's important for us as a society. Yes, I hear the the phrase of AI slop. What I really hear is that there's low quality low quality contributions at scale. Okay, we also used to call that spam. What do we do about spam? So I think in terms of like really breaking down and taking away some of the heat and some of the emotion around what the current kind of either solutionism or the current productization is, and bringing it back to problem spaces and asking whether or not this is even the problem space or the solution space we should be working in, but really driving that down into like I told you, it's no fun at parties. But driving that down into a place that we can have conversations around complex areas that require a lot of work and ideas and collaboration.

Ilan Strauss  57:52  
Amazing! Thanks to our four amazing panelists. Thank you so much, and thank you to the audience for engaging and listening.

Announcer  58:02  
The Engelberg Center Live Podcast is a production of the Engelberg Center on Innovation, Law, and Policy at NYU Law, and is released under a Creative Commons Attribution 4.0 International License. Our theme music is by Jessica Batke and is licensed under a Creative Commons Attribution 4.0 International license.