
How SPUR is building a new economic model for publishers - with technical lead Alex Springer
About this episode
Publishers are not sitting idle as AI companies scrape their websites and cite their content in response to queries without permission. As their business models are turned upside down by sharply declining referral traffic from search, they are fighting back. That starts with understanding precisely how AI companies are scraping their content.
In February, a coalition of leading UK news publishers founded the Standards for Publisher Usage Rights, or SPUR, an effort to develop and formalize a way for AI companies to make use of journalism in a way that protects publishers’ intellectual property.
Dubbed “NATO for news” by FT CEO Jon Slade, the coalition has since rapidly expanded beyond its original founding members, and in June it released what’s known as a draft telemetry standard, an effort aimed at providing the technical foundation and shared language necessary to understand how AI systems use publisher content. With this knowledge, publishers can then start to create informed licensing deals with AI companies.
Alex Springer is the technical lead at the SPUR Coalition.
He spoke with The Media Leader about SPUR’s work, and whether publishers have finally developed the leverage needed to get AI companies to pay fair value for the work they’ve been allegedly stealing.
Highlights:
4:10: Redefining publisher value in the AI era
11:11: The SPUR Coalition: Goals, technical developments, and building co-equal relationships with AI companies
20:01: Can the long tail of the web be protected from AI copyright theft?
26:41: Is the publishing industry being incentivised to improve quality?
35:22: SPUR's draft telemetry standard: what it is and how it works
43:04: What a new marketplace between publishers and AI companies could look like
51:44: Government interventions and development timeline
Related articles:
Telemetry: the key to solving AI’s valuation and measurement crisis
How publishers are fighting to protect their copyright from AI companies
Get every episode summarized
Each time The Media Leader Podcast publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Transcript ready
667 searchable segments. Every word is indexed and playable.
Full transcript
The Media Leader Podcast — How SPUR is building a new economic model for publishers - with technical lead Alex Springer. Machine-transcribed; use the interactive transcript above to jump the player to any line.
This episode of the MediaLiter podcast was edited by our production partner, Trisonic. And if you're looking for podcast production support, we highly recommend them, find them at trisonic.co.uk. We've just decided that because this content owner side of the world is so fractured, because there's a billion of us and it's hard to organize a billion humans to create content, that that's the one we can exploit. We have to find a way to actually talk about that. Welcome to the MediaLiter podcast. I'm Jack Benjamin. Publishers are not sitting idle, as AI companies scrape their websites and cite their content in response to queries without permission. As publishers' business models are turned upside down by sharply declining referral traffic from search, they are fighting back. In February, a coalition of leading UK news publishers founded the standards for publisher usage rights or spur, an effort to develop and formalize away for AI companies to make use of journalism in a way that protects publishers' intellectual
property. Dubbed NATO for News by FTCO John Slade, the coalition has since rapidly expanded beyond its original founding members, and in June, it released what's known as a draft telemetry standard, an effort aimed at providing the technical foundation and shared language necessary to understand how AI systems use publisher content. With this knowledge, publishers can then start to create informed licensing deals with AI companies. Alex Springer is the technical lead at the spur coalition. I spoke with Springer about Spurs' work, and whether publishers have finally developed the leverage needed to get AI companies to pay fair value for the work they've been allegedly stealing. This is our conversation. I hope you enjoy. Alex Springer, welcome to the show. Thank you, sir. You had an interesting career and journey into working now with publishers, and I want to get into your current work, but first, how did you get into the media industry in the first place, and more specifically,
working with publishers to measure how AI uses inside their content? It's still fairly new to me to even be saying I'm in the media industry. Actually, I would check perspective I was working in marketing tech for the last 11-12 years, so measuring, reporting, tracking on user clicks and interaction with publisher sites, but really from a brand and advertiser perspective and an understanding user behavior. Big focus on the affiliate and the influencer space, and that kind of led its way into the existential threat and the existential question that's now happened with AI, this says, well, if users aren't going to visit your content anymore, what does it matter? We can all know it matters, like we see it being cited, we know they have to be trained on this stuff, but particularly in the affiliate and influencer performance tracking space, if there's no click, then there isn't really an industry. The question was, if that's the case, there's no more clicks, we gain that all the way out and what's the new thing? How do we define that new thing?
And a lot of the telemetry work that we'll talk about in the reporting and the content usage language that we started defining came from that space, saying, okay, what's the new unit of value? It's no longer a click, it's no longer an impression, which is not totally true. To be clear, there are still clicks and impressions happening, right? But in AI systems, no longer a click, no longer the impression, what matters is it a citation, it's a sort of thing. And so that's led me into the media space working with SPIRR, which I'm sure we'll talk quite a bit about, but I also think of it more broadly as a site owner and just general knowledge problem. So I will happily work with brands on their product feeds as much as I work with the Guardian on the articles that they have, because from AI system, we're all content owners. And that idea is actually really, it's quite fun as a unifying thing, because you get to bring these people together that are non-normally interacting. So I'm on the media side with my head on for SPIRR right now, but then I'll leave here and go speak to a couple of brands about how they can make sure that their intellectual property is correctly represented, right? And they
get that visibility. It's sort of a redefinition of value for this new AI-led era, a little bit. It used to be circulation was probably how most newspapers would have measured and sold their advertising based off of, and then it became search and referral track, and click and that's why you saw some SEO practices crop up, but then also click that type of stuff. And so now we're moving into this new era of, well, how do you actually publishers get remunerated forward the work that they're doing? Yeah, absolutely. And that's a good point too, because that transition also is something that represents something pretty significant, right? So circulation-based, subscription-based, has a certain thing you have to do to get that to continue happening, right? You have to deliver value, you have to have some sort of authority, there has to be some trust there. Ads change that pretty drastically, right? Like I no longer really need you to trust me and subscribe and this kind of thing. It's an option, and it's worked for many great media sources. I need your eyeballs,
right? I need your attention. You've for a second, for one second and 50% of the pixels that I need to show you so I can get paid, right? Like a map fundamentally changed the equation, and so subscription started going away. Circulation maybe started, it's slightly less than it was more about audiences and that sort of thing. That has now come back to light us in a pretty significant way, because from the outside, it looks like the internet is free, and the AI agents all showed up and they described everything, and they said, oh, well, it's free. It's all there. It's great, right? But really, what was happening was we were saying, well, it's not for you're paying with your attention, you're paying with your eyeballs. And if you strip out that part, that payment portion, you've now rendered it free, but like removing the monetization from something and taking it is called theft. And so we need to now move back into the scenario of saying, if it was authority originally, that would drive subscriptions, and then it was attention and grabby-ness, shall we say, that was the drive and the impression. Sensationalism, yeah. Sensationalism, yeah, sure. See, this is how you know I'm new in media, making up words.
No, I like attention and grabby-ness, I think that's a new coinage. And now we have to go, okay, we could either stay attention and grabby, and go, I'm going to grab the AI agents' attention. And so to do that, you have all these SEO, turn GEO practices, right? That sucks, right? Like, we can have AI overview, and you see, oh, here are the six top quoted articles, and you click one, get to that website, see a flash, and then you're suddenly on the homepage of that publisher, and that website, the like, overview sorted, but if cited, doesn't exist, right? Or is it never meant for you to see? It was entirely game for grabbing the attention of this. Is that commonly happened? Oh, yeah, absolutely, right? And that's, that is unoption, right? We can be attention-grabby-notention-grabby towards the AI systems, and we can do this crap. And then not only is that, create a poor experience, and it's going to get wiped out, right? Like, that's, Google is very familiar with people trying to hack its algorithm. You know, it might not be tomorrow, but it will be very soon that they make that no longer something that happens. The opening I did with Reddit, what, like, a week ago, right?
Yeah, I've seen those reports. Yeah, and so that's the attention-grabby option, short term thinking, applying this old model to this new thing, right? Or we can think about how we evolve, wandering off into the rant a little bit, and we'll get to some of the other questions, but I think that finding that new economic model is what becomes really critical. And I'm hoping the work we're doing at Spur and the broader reporting work that we're doing is laying the groundwork. It's not the end goal, but it's the groundwork to create the sort of communication and transparency we need to start defining that new economic model. Do you think an ad-funded model still will work in this future, or subscription-based models will become effectively necessary sort of table stakes for publishers if they want to properly monetize, sustainably monetize their businesses? Yeah, they're really a blend, for sure. And a little bit is you have to ask, like, what's the who's paying for it? Who are you providing
value to? So if it's ad-supported, then there's still traffic going to websites, right? It's dropping pretty drastically, but there are niche websites. There's reasons you go to a website, it's a seven AI agent, right? Like we like the experience of the news, that kind of thing. So some of this is a little bit contained into like commodity-type information, right? Things that I just want the facts that don't want to have to go read your entire recipe blog just to get the ingredients to your msue, right? Like, like, those sites can go away, and I think we'd all be pretty happy, right? SEO sites. So as that were made to sort of game Google can go away. Exactly, right? And so, but you look at things like economists, things like Atlantic, right? Stuff that's packaged together as an experience, whereas a human-going game information, not an answer to a question I already know what to ask, but instead, I learned new stuff, right? I learned new things to ask. Maybe I started with one kernel of a question right by the end, I'm richer for the experience, right? I think those sort of things bundled together get to be really interesting. Not because they're necessarily sold to AI systems in that bundle, but because
they're still, they have a very viable model, whether I subscribe or whether it's ads or whatever it is, I still go visit. That's not my sense. The subscription inside of the LOM starts to get really interesting, right? So if I have a subscription to, I'm going to keep picking on the economist. Do they necessarily want to unbundle that and say, yes, now your cloud can read your edition of the economist on your behalf and disarmediate the economist? Probably not, right? And everyone's experimented with everything, right? But like my guess would be that that starts to devalue a core proposition that is the experience that is the entire bundle, right? Is there some version where a publisher can say, well, if you look at all of my content, I'm an expert in travel to Mauritius. I could either let it be scraped and make the AI agent and try to figure out what's important about Mauritius based on all of my content, which is really inefficient.
Or I could start thinking about developing these skills or an agent of my own or some layer that interacts that distills all of the information I have about Mauritius and then says to the agent, oh, you want to know? Right? Tell me what you want to know. Ask me a question. Or here's my skill. I'll make you an expert. And maybe there's some subscription there, right? That says, I as an agent, don't subscribe to the entire encyclopedia. I subscribe to a really tailored high quality skill. That makes my agent smarter, right? And the publication, the media folks say, great, I earn my living off the subscription, but I'm making sure my skill is really up to date. And I do that through investing in journalism, through research, right? Through all the stuff that we do today. I'm not saying that's definitely the model, right? But there's lots of options to innovate. We'll talk a little bit more about potential business models and neighborhoods a little bit. But I do want to talk about spur. The spur coalition was founded in February. It's what FTCO John Slade called the NATO for News. Your eyes just went blind. Like you can't believe it's only been
eight months or so. I think that Slade evoking a military alliance seemed fitting, given the publishing industry, is in a real battle with AI companies that are scraping and using content from them without permission most of the time. For those unfamiliar with the project, what are spur's aims and what are you working on as the technical lead? So spur's overarching goal is to shape and guide and help usher in this economy that works around information. The other founders of spur will say that's like differently, right? I think the NATO for News figure is great for headlines, right? And there is some defense stuff that we have to do because an incurred state. Yes, the lot of stuff has been taken. A lot of content has been repurposed. There hasn't necessarily been remuneration for that content going back to publishers and that's a problem and these defense. But obviously the overarching goal is to find this future where
things are viable, right? Where the AI systems can get what they need, which is good inputs so that they can produce good outputs. And so they can do so without devouring the system that creates those inputs from the spur side of things, the publisher side of things, right? We want to make sure that we have an awareness of how our content is being used. We want to know what content is needed. So like any marketplace, we say, you demand folks, what do you need? How do we create the supply for what you need? Right? It doesn't really matter if it's an AI agent or human readers or whatever. A functioning marketplace requires some kind of communication back and forth so that the supply side can create what the demand side needs. Because we're talking about like knowledge as the product here and information as the product here, that communication gets even more critical when you start to think about how is it reshaped on the way out and over, right? The part of the value of information, particularly in this this world where you can get the same news from 50 different sites comes from the authority of the
supply source as well, right? And so a big part of what Spurs is looking to do is make sure that we create an environment where AI agents can go, get what they need, know how to pay for what they need, to make sure it's licensed, ideally in a way that makes it safe for them to operate. It's more efficient for them to operate, right? They don't have to like reprocess the same 7,000 HTML tags off of some crappy script thing. They can communicate back and say, oh, hey, this was not quite yet. I need something a little bit different, right? That's great from the advanced side. Supply side, and you can confidently start to do things like make businesses out of the information industry, out of journalism, out of content creation, because you have those signals because you know the demand is there, right? Because you can engage with it in a productive way. My hope is that we get over this hump at the moment that is really antagonistic, and we find a way to reset a little bit. Because if we were to start this thing tomorrow, right? We said, okay, we have these AI systems, they happen to exist. They didn't already steal everyone's stuff. We're not suing everybody for everything all over the place, but they came up and said,
hey, look, we really need a whole lot of content, right? And we are well-funded, and we don't want to risk getting sued. And so we're going to start the relationship by saying, you all create great content. You spend centuries figuring out that are right words that make a lot of sense and turn the world into something sane. Let's buy it from you. So we're trying to build that marketplace. You mentioned that AI companies are well-funded. I think it was Nick Klag earlier this year who said something like, actually, we wouldn't be able to, we, sorry, the AI companies wouldn't be able to actually pay for all the content that they have allegedly stolen from publishers. And we've also done some reporting on the financial pressures facing a company like OpenAI. And that's why they're getting to advertising. So are they actually well-funded enough to cut deals with publishers that give fair value back to the publishers that are creating the content? Yeah. And there's a difference between funded and profitable.
I mean, there's two things to that. So one, if you can't afford it, then you don't get to use it. Like, flat out, like, there's no, that seems pretty straightforward. Yeah. Why has that changed? Again, it's this conception that the internet was free, right? And then so I get how the more naive amongst us may have gotten there. But that's just basic tenets. Maybe we've outstripped. What we can't afford it from the resource we have. And then we have to reassess that and forget that. But like, there are other things that just cash. Right? And they know this well, right? I'm not saying we go give stock options in OpenAI to every publisher, but maybe, right? Like, SpaceX goes public and all of a sudden, all the content sources that GROC has used to then be part of that revenue share in the winnings. The same way that was the janitor that's now a millionaire kind of thing, right? I don't know if you want SpaceX stock, but you know, I want to say enough. I mean, there was a moment. There was a moment. There was a moment. The one big thought economics didn't matter anymore. But yeah, I think that's an interesting idea.
There's no, again, I'm not saying it's the right one, right? But we're missing that room where we can sit down and say, look, how do we make this work? Right? Yeah, sure. Okay. If we think that it's important to be able to train these, I've ever ever more powerful systems. And we're willing to pour money into compute and resource and power and water and all these things that are also pretty limited resources at the moment. Because this is so important, we've just decided that because this content owner side of the world is so fractured because there's a billion of us and it's hard to organize eight billion humans who create content that that's the one we can exploit. We have to find a way to actually talk that out. Why did it take so long to get to spur do you think? I mean, when Chatsch.ubt launched in 2023, the very first concern I had and I am not special. I think many people have this exact same concern is no one's going to click on links anymore because the information has just served you and that would undermine the business model of basically the
entire internet. So why did it take two and a half years to get to the point where publishers are actually coming together and forming a coalition and saying we need to develop a way of tracking what's going on to our website so that we can create licensing agreements to our benefit. That activity has been been happening since 2023, right? Like that's that there are a lot of really good efforts that have been around almost from day one that have identified that issue and are trying to do something about it. You've got a lot of in addition to like the bad actors of the scraping world around the exos and the taviles and these folks, right? You've also got for everyone of those, you have someone like a told bit or pro rata or cashmere or red pine. These folks that pretty immediately identified let's figure out a way to serve publishers and create a market for this stuff. Really simple licensing and which is RSL standard and RSL collective
are doing great work around machine readable licenses and they've been doing that long since before. Spur started. Then individual publishers are innovating really fast in some instances. So the folks that people in in Washington Post and like there are a lot of different versions of things that are appearing which is great. Spur is a coalition. Took a while to get together and I can't take credit for this. This is a, this is a guy named David Buttle who lost a lot of sleep over this. It took shape in order to have one really explicit purpose, right? So look at this wide range of efforts that have all kind of sprung up and start to do a little bit of organizing, right? We still think that like the market will find a solution to this and we very much want to lean on the market. Do so. We're just trying to provide a room and a conduit to communicate across. There's also just an element of getting to the point of eight billion content creators running around the world, right? It's just tough to organize competitors and making sure we're doing it right is really important and ensuring that we create a coalition of content creators
of journalists, ogadia that can approach this in a way that is legally compliant but is also something that is beneficial to the industry rather than just to a few. It was founded by initially a few. You had five first founders as Guardian, FT, Telegraph, BBC Sky and that happened in February and then there were some two dozen members since then that have joined from several different markets. It's not just the UK anymore. That said, I don't necessarily see a ton of small mid-market independent publishing brands that are necessarily a part of spur at the moment and those are publishers that I think are really at the coal face of the changes happening in the online ecosystem. I've talked to independent publishers who have had to go under or know people that have been had to lay off most of their editorial staff and outsource all their content creation to people and other markets because they can't afford to pay the labor anymore and there's not necessarily a
whole lot of certainty of well let's hope we will get some sort of information about from spur maybe or some framework by which we can create a market. They might be gone by the time the market is viable for them as opposed to I think I think we decree the FT could probably last a couple quarters of some tough business and be okay on the other end. So how do you think about the long tail of the internet and how what you're doing could affect them if you move fast enough I suppose? Yeah. The authority issue we spoke about earlier kind of comes in play here a little bit right? So you think about the differences between some of those big publishers and the small long-tail ones. Part of it is the breadth of content right now it's a lot easier to just go search FT wants to get global news then to maybe kind of sift through all of the different smaller new sources that you might need. So there's this convenience thing that hasn't quite been solved yet.
Some of the scraper's and the grounding services that are not paying publishers have solved pretty well right they keep using this phrase that they're organizing human and his knowledge like they're doing us a favor which you know getting paid for that knowledge along the way would be useful but like so the of the access issue how do you get the thing you need very quickly and spreading across all these sources is difficult not to say like what was me on the AI platform right we have to solve that. There's also an element of authority so if you get 20 sources that say a thing and you're a charity me to you're trying to convince the user to trust you and it turns out that a really easy way to do that is to err on the side of quoting a brand that you think they'll probably not know and that ends up defaulting to things like FTE, PC, New York Times right. They also quote a lot of like small blogs that they think are experts though. Right and this is where and so now yeah you do things like so one great way to stay stick around at the moment as a small publisher is find that niche and then like defend the hell out
of it right so you can keep getting quoted on small blogs that also then say if you think the order the order of operations says user asks the query I did a web search of some kind based on whatever service I have I gathered back the results I have a layer that says knock out results that'll get me sued right because everyone's monitoring right now so Chatchock T says something like oh the BBC said something and here's a BBC link right red flags go up because there's between robust RTX D and blocking mechanism and stuff like that's not supposed to be happening. Okay so now you end up with what's safe as well right and so you end up with these struggling smaller publishers that are making decisions like with two days worth of revenue left trying to figure out okay do I block do I unblock like I'm losing on my traffic am I shooting myself in the foot by doing this what choice do I have because even if they get cited and and you know their website comes up to a user the user is still probably not clicking through to their website anyway so there's no value that's been created that can be passed back to the publisher so you kind of in this damn if you do because you can't
wait right to your point like you can't sit and wait because you don't have that war chest to be able to do sorry say well okay even the trickle is better than nothing right or or maybe maybe I can turn that into like an AI authority signal because all the affiliate performance folks are saying that like if I can get your brand mentioned they'll pay me like per post now right like maybe there's the other things I can do and so you end up with always bitty little sites popping up as as references and citations in systems like shaggy bt because they've tried the authority route they'd been blocked or they don't have a license whatever it is and they're trying to avoid giving more evidence towards towards lawsuits so then they go often to try to find these kind of secondary authorities or these other smaller things right and put that all together the system itself is much more complex and that right but like that's the sketches of what's what's happening there that's part of why it is so hard to figure out like how you optimize for things like AI answer engines and how you get mentioned is that black boxes has complicated more complicated than pay drain can all these other
algorithms so that that's the pressures that are actually kind of like hitting the smaller ones first right and we're not going to spurs not going to have an answer for that tomorrow or expo whatever it is I think until we have some viable route to a marketplace AI systems are substitutional products that are not sending revenue back right like there's no and I'm happy to be proven wrong but like I've not spoken to anybody or seen a model where leaving your content open to scrapes gets you anything back no matter what size you are at this stage right you block you potentially get a license deal if you big enough if you have the content they need if you don't you've lost nothing like it's a there there's no gain there right and then you find other avenues to find your audience to connect with that audience and then you hold on working with stuff like spur and offering support there right so we do have four up to over a hundred members now so the founding group is the big folks for that at AP and
and CMA and media host to that group but then we've also kind of running the spectrum on small publishers joining as well as well as some affiliate memberships for like tech and intermediary platforms and that sort of thing so I invite folks to engage with us in that regard right we can also link you up with the intermediary tech platforms that are working for publishers and trying to find solutions right and we're also working with folks like the NMA and NLA and these other member orgs that have been around a lot longer to try to make sure that we're not duplicating efforts right and that we're able to kind of do what we can do as a new kids on the block but also then lean on the strengths of people who've been working for publishers for a very long time. Something in terms of publisher strategies that I've started to notice is that they're shifting from a sort of volume based approach to a I mean I would argue sort of quality based approach essentially not caring as much about traffic not sort of deciding we don't want to care exactly how we're being cited or if we're being cited because as you mentioned it's not necessarily
doing for us anyway and traditional search also isn't doing for us like it used to so let's actually focus on publishing really great original journalism that we think has high value building a readership that way monetizing them through membership subscription what have you this is maybe giving the AI companies way more credit than they deserve but is there an argument you could make that this change is sort of incentivizing the publishing industry to get better at what it does. In some regards I think it I wouldn't give the AI companies any credit for this this has been happening long before tragedy T right like this this crisis of search creating duplicate products right people looking at snippets and not clicking through right like snippets were there before AI overview was there this is this is not an even new issue I was a Ken lion's a month or two ago I was and such a blur I mean you said February for Spur right and
there's this thing I kept having this kind of like vaguely uncomfortable conversation with people around like this is partly us being found wanting right it's not AI that is necessarily driving all of this but there are a lot of users like I came up to you I say okay I'm you have the internet it's shiny it's pretty it's covered in images and flashing lights and all the information you could ever want and it has a really really good way to find all of it like Google search is pretty damn effective at getting you to the page that has the information that you want that's door number one door number two I have a text box it's ugly it lies half the time no idea where the content came from in the first place most of the time no one really likes like it it's obnoxious it's flattery it's sycophantic but doesn't have any ads doesn't have much for now for now doesn't have any flashing lights it claims to I mean they were Google ads running the cinema like ahead of of feature films that were like just showing people being like I'm
gonna find a dessert in Italy oh that blog is too long read it for me right in some ways that's the users and consumers fighting us wanting and telling us that they'd rather deal with the fact like a terminal like a text editor buck thing then go deal with your and should have fight internet right and so in some ways that is going to be driving the people who create information and knowledge to do better in getting relevant important stuff to be focused on where is there maybe a new angle or what can I add to the conversation or what's different right um and I think that's a good thing it raises really uncomfortable questions about how many different like loudspeakers we need amplifying the same piece of news um and how many the economy can support you think there's enough supply basically of news publishers yeah and maybe right in in some areas absolutely uh depends on what news we're talking about right I
the proliferation of like outrage based news probably created too many beacons of outrage and what do people do when they get stressed like they avoid the stressor right except for the ones that really lean in and like that became the thing we stratified the two like you think a really outraged and consumed consume consume or get so turned off that you go talk to a sick fan that lying in chapeau right so I'll have to say I think yes this is pushing publishers to find new business models hopefully in some ways to communicate more effectively um I think the winners that shake out of this or maybe this is just hope but I think the winners that shake out of this are the ones that create effective important information that help people do something better right whether that's be entertained or know what's happening in the world or just be educated in general hopefully the other option is that we all end up as just signal producers right and like everything's listening all the time goes up into a big morass of information and that gets mined by open AI uh and like we did paid pennies for our news which that kind of sucks what do you
think of a published like the New York Times uh working on developing testing its own version of an AI chat bot and sort of saying and I've seen this the same demonstration from other publishers as well I know it's not just NYT uh uh doing this but that was a recent report on it from semifor sort of that habit formation of I want to ask a chat bot something as opposed to a traditional search engine yeah well as a consumer maybe I don't trust the entire internet that is feeding into a chat GVT but I do trust the New York Times journalism and I'm happy to do that on a specific publisher's website yeah you think that's also a potential model absolutely yeah and I get back to some of that kind of like kind of still all of my content into a skill that you can interact with right sort of thing um yeah 100 percent so my my biggest fear around all of this is that same home and has this interview somewhere where he uh he like has this little chuckle that he does he's like oh I know it's really weird man like we're all we're all like talking to the same brain like I just I wanted to shake this framework it's you just it's your
company like you like that's terrifying right um we've created these kind of general purpose large language models that are very good at producing language uh and then in order to make them acceptable or safe according to whatever standards or whatever culture that they're deployed in we've tuned the hell out of them right to agree with the cultural norms of where they came from to agree with the opinions of whatever is like the correct thing in that society in that culture and in some ways that's really really really good right like we don't want our chat bots to be racist and misogynistic and these sort of things but there are plenty of studies i'm plenty of looks at how the cultural norms particularly of something like silica and valley where they're all generally trained and tuned start to really come through strong right like the white male tech bro culture out there filters through and that's not just a demographic based thing that is also this kind of big warning sign to me that if we cannot get transparency into how these things are tuned trained and deployed then we shouldn't be by default trusting it particularly with things like
our entire information ecosystem that ran to side the deployment of an agent by an expert in a space would seem to be a good attitude to that right um we still have issues with the underlying model coming from certain sources right so like New York Times agent don't know what model they're using potentially let's say it's running on chat to be tea right so as long as GVT 5.6 it's going to inherit all those biases you know and we're still beholden to a group of faceless engineers that we can't interrogate about what they told that model to start thinking about things but to very least now they might not even know i mean i think that's what's what's what's stunning absolutely right and um and so okay so we're beholden that but then we say right New York Times is going to take that we're going to build an agent on top into your point i trust them about their reporting more right or or insert media company here right like don't really care where your your your biases your media source lies here it's it's it's i want to know telegraph trained agent because i trust where that's coming from i want to know New York Times trained agents i trust where that's coming from
i want to know Cleveland Clinic Health System trained agent because asking Claude who is increasingly a coding expert to talk about i don't know like liver issues it doesn't make a sense for it um and so these specialized models owned by the actual knowledge creators tuned and trained by them i think is a really exciting feature um and it's a really exciting application of this technology uh there's absolutely nothing that says this has to get concentrated in the hands of four giant us based firms um we've just like pretended like that's been the default because that has to happen right and yeah there was a lot of capital and a lot of resource that went into it but like we have open weight models and open source models um we have specialized models we have a lot of innovation happening in this space um and while open AI and google and anthropicon folks are are struggling to retain their hold on this idea of like one platform to rule them all i think we as consumers and we as an industry should look at this alternative where we can democratize this and we can give our experts tools to amplify their expertise so in order for publishers to
to amplify their expertise and to create licensing deals potentially with AI companies as well in in order to create sustainable business models um they need to understand what's going on and how AI is actually interacting AI agents and platforms are interacting with their content uh which just sort of takes us to the draft telemetry standard that you released in June first of all i you know i appreciate you the technical lead um telemetry was not a word i was familiar with i think until i read this report so in in layman's what is the telemetry standard that you're working on what's it trying to accomplish why is it called telemetry yeah um kind of kicking uh it's one way it may not be uh so so telemetry at core is just an idea of a closed system signaling out what's the current state or what's happening inside of that system right um and that telemetry signal can be picked up by anything that is listened or anything that is listening anything that can be authenticated right and so you can then kind of start to measure what's
happening inside of the system um and you'll see that around uh most things right like what is the position of this widget inside of the machine at this stage um can send out a signal but we're looking at specifically though is telemetry around content usage events um so this could just as easily be called a content usage vocabulary as it could be telemetry um which might have been a better way uh no i think it sounds really cool actually i think i think keep telemetry there's a phone home i don't know yeah um and so we want to do though like everything else aside that idea of what is the unit of value in this new ecosystem right it's not an impression uh it's not a click um from a content perspective it's did the content influence the answer there's been then shown to the user right was it used a time of inference which is at the time that the query happened that the chat is actually happening um you know order to make that answer better in some way to shape it in some way that's sort of like right uh which is to say did that content deliver
value in the course of the agent trying to deliver value to its user and as a proxy for that uh uh right we say okay great let's look at the events so um we start with this chain it says the content itself is retrieved uh that's a website scrape that's being pulled from an API that's being pulled from repository but the the content has now crossed the boundary and it's now inside of the system that the user running many many services one of which is an LM right in order to start to generate a response for a user um that scrape event is where a lot of the industry started trying to figure out it to be charged per scrape right to charge per access that model is not currently shaking out very well why um partly there's a like a scaled issue that says okay well like the pennies per scrape and also just lots of scrapes happening and then also it wants to scrape it once do you just cash it and just keep it and now i have it myself but i don't have to go scrape it anymore like everyone charged me per scrape why would i come back right um and then like license
terms say you have to and all that stuff but but also the access itself doesn't show value right access itself says that i looked at it um but maybe it's not what i needed you know uh and there's an argument that like if you walk into my shop and you picked up the snickers bar and you walk out of my shop then you pay for the snickers i don't care if you eat it right you pay for the snickers bar or you just looked at the snickers bar in the shop and you didn't do anything with it and you bought something else right you could also make that argument and that instance exactly that's so so we've then said okay the next stage so content retrieved as an event the matters we want another scrape we want the access we also want to know that and was it grounded right and what this is is say it is the content has it entered the context window of the agent right so when you send a block of text into a large language model um you that model then operates on that query and then returns a response to based on what's in its training weights and what it's been trained on but then also what's gone into that context window um and that context window you'll hear about like the ever larger context windows of the next greatest model that has a million tokens and all this stuff that starts
to matter because that's where you can really influence the output of the model right so if you were to pop the hood um look under the bonnet and go say i have a centaquarian it says what is the weather in London right and AI overview doesn't just send that to Gemini right like google will say okay cool let's do some search results great um we got 10 results five of them are for London Ontario five of them are for London UK google knows that you're sat in London UK so it throws away the five London Ontario events right so it's retrieved all 10 but five of them are not relevant and so there's no reason to put it into this valuable real estate that is the context where there's no reason to send it to the large language model to pay for the tokens to pay for the inference that is required um because you could just say oh here model here's 10 things you sort it out right but instead we can do this pre-processing um and it's a lot cheaper and it's a lot more deterministic and it gives better results throw away the five and we see great now we have the five London UK weather events now let's send that into the model right so now we have your query what is the weather in London and a little pending that you'll never see that says based on these results
Gemini what do you think the weather is in London or answer the user right so Gemini's not looked at the window right Gemini's looking at these five results and that's been so now you have the five grounded results that have come and see if your content graded right from there we say okay was the content cited to the user so once the generation happened once the output happens do you see the links back to this um and thinking like traditional bibliography terms right is the user able to go get the content from its original source right and that's our citation event um and then we've got a couple on there around things like was the content displayed in its original form and then did the user click it engage if we can get that vocabulary baked in as default and to be clear like this is already happening like there if you build an AI agent you're monitoring what goes into the context window and coming back out as part of your testing process your e-vail process right this is not like a technical leap to do this what we're trying to do is say can we bake this language in across the industry so that's helpful for reporting back the value that content delivered because you can say well like I retrieved 50 things off your
website only two of them made it past my screen process I'm not going to pay for the 48 but like maybe I'll pay a lot for those two right because those are really important um and then you look at the output and you say okay to the agent help the user in the way that we wanted right if it's chat's BT did it lock the user into an endless conversation and isolate them from their friends great high value we want more of that content right um if it's a customer support agent and it's like oh did the user come in and have one question get the answer they need to get out right because that's our measurement of measurement of value um and so then we have this getting back to that communication thing where the demand side can say based on the events here and I can communicate in a way that you understand because we have a common language uh the odd and grounding all the time on your stuff and as a result it's much better output from my agent right the user gets better value now I can justify paying for your content um if we're missing these reporting points in this conversation that one's really hard to do that it's kind of just saying well uh and this may I may not be what chat's BT just do with Reddit right they said okay I don't know for sure this
is that one but like they're really negotiating a deal right now and if they're saying hey turns out if we just turn you off our users are still pretty happy great I'm going to pay you half what I paid last time right because I there's no value here right if we had the ability to actually measure grounded events tied to citations tied to those outcomes um then you can start to have a real conversation about the value of that content is so you mentioned a cost per scrape model not working so far but you have sort of identified five different steps that a model takes where do you see the marketplace maybe ending up I mean I imagine it's still in flux but yeah generally speaking our publishers the publishers be most interested in uh trying to get a cost per citation versus a cost per scrape versus a cost per I imagine not clicked because again click through rates are way down so it's important to know all this information yeah but then you need to make a decision on where do you think the value lies yeah coming from the brand side there's this pretty simple little like tier daughter seesaw you can see right where uh I think ad tech land you have publishers
brands on either side of the seesaw right the bigger the publisher is and the smaller the brand is the more likely you're gonna end up with a kind of impression based relationship right because the publisher can say well like all I really control is gay eyeballs to my site right but I'm big I get a lot of eyeballs and you're a small brand so like I can send you traffic but like whatever you do with it doesn't really matter to me right I'm not gonna tie my fortunes to what you small brand does um right reverse that you've got tiny publisher you've got someone like an apple or something that's saying well we're not gonna pay impressions right like I don't I don't it's not worth my time you have so little impressions I'll list it but if you send me traffic it converts then like I'll pay you a commission right and so on one side you've got this kind of cpm impression based and then cpc somewhere in the middle and then you've got your acquisition right and that power differential starts to define who gets paid on what events and and that holds fairly true across across those industries I think we end up in something similar here all right so like a per retrieval first rate that's that'll be a viable model somewhere um you'll get paid a lot less first rate right
whereas like an actual per citation probably ends up having a higher fee per citation um per engaged per click maybe even higher still so there's an element of that and going okay well where it's it's just more negotiating levers based on our relation there's also a difference per industry right so if you think about like 11 labs does a really good job of actually engaging with content creators right and they're a voice generation audio generation um AI system uh it doesn't make a sense for them to have citation events right they they ground with audio that they've licensed and then they create voice prints and audio and user doesn't make it like you don't have a tag at the end it says oh this was also sourced from 17 different people or whatever right um and so they're gonna just pay on grounded events because that's what they do a day that works so well um there is a distinction between the telemetry standard and a telemetry profile inside of spur it's really important so the telemetry standard is always spoken about it's the vocabulary it's the words it's the definitions of the events themselves there's no opinion in there about how to use those
and that's by design the spur shall emetry profile is very um media specific right it's saying okay if you're gonna be compliant with the profile you're gonna send us back data about the events in the standard format but then you're also uh going to do to a designated endpoint the publisher chooses so you send that data to me um as opposed to keep it inside your wall garden um and you're gonna send us all those events rather than just picking and choosing which one you're gonna send um and they should come out in real time or near real time uh so we're not getting like a monthly aggregate um and that distinction is important uh the left bit on kind of that standard thing the the thing we really want to make sure is it's not just media not just publisher so i think that's where it goes to end up spinning right you talk to security teams you say hey what do you think about content entering your agent right i say i think that's a prompt injection security nightmare right and select those grounded events are really important because i got to look at what's coming in the front door um and if i can get premium content from the sanity like from people
linked that is sanitized and i know it's safe um that's one less worry i have that it's going to be content is randomly scraped from a website that is causing my agent to dump out my passwords to someone in turn is up right um and so we can get this really cross cutting buy into the vocabulary where it's useful for security teams it's useful for publishers it's useful for AI agents and my hope is that we're not even talking about it in a couple of years it just is the words that we use right um and then we can all move on to creating dashboards revenue and all that stuff around that and the market can be flexible as flexible needs to be any individual publisher brand whoever can cut whatever deal that they think is it is reasonable that companies agree to do you think that publishers brands have the leverage over AI companies at the present to bring them to the table you know no so at what point does that leverage start to actually exist yeah um i mean i think there's there's some right uh and you know like folks like
LeMonde is a good example of like right there they're doing a lot of good work there around kind of making their own path um and like not lots per members right and um very kind of vocally out to do what's best for LeMonde um and they have a unique data source right i get they're in their language in their market they're the heavyweight and therefore they got that leverage i mean publishers broadly struggle to get that leverage because there's so many different ways of getting content information and so yeah the big ones right if you're the guardian and ft in LeMonde you can sign license deals um a lot of those original deals are like payments for damages done basically um there was no real pinned like what is the value delivered into the system of this content um and because we're missing a language to describe that value as they come up for a new way you start to really struggle you start to say what's this new value for with when an AI company is going to review uh and then renew presumably some of these licensing deals uh why would they need to if they within the original licensing deal yeah they would have
gotten the entire library of that publisher's content presumably so anything new that they pay for would necessarily be of much smaller value because you're only covering the amount of time between you know the last time you inked the deal and the current time that you are inking a deal yes i would think puts publishers in a really difficult spot it does and it depends i can't speak to individual deals right but it's it depends on how the deal is into the first place right because there is an element of saying like yeah in perpetuity versus you know do you have to purge your all of our stuff out of your records and i think um and it depends on how that was negotiated and who who did the work and good luck getting compliance if that is the case good luck getting compliance especially if you don't have some kind of measurement language that you're getting compliant against right which is a big part of what we're trying to do i mean that that is the question right like the iceberg of human knowledge is there taken and the stuff that's above the surface they were creating on a yearly basis that needs to get renewed into is increasingly shrinking and that's that's the challenge um it's part of why we focused on the inference time
as opposed to training with this burr work there's a lot of people trying to solve that sort of kind of training bulk data problem um and doing some really good work there but that's sort of recency thing gets really important i think it points back towards this idea that we need to reform the market around knowledge rather than just try to take the old principles and apply it right because you end up with this like we're here to set a conversation so iatf is doing some really good work on a standard side right now uh trying to give publishers the tools to express their preferences on how they what they let AI systems do right um and the argument that the moment starts ending up being well okay if a user copies and pastes stuff in to chat we see and says go can we even do we have a preference around that like our preferences maybe that doesn't happen right but like the user you've given the user access and i just do whatever they want right and we've never really had this great set of controls downstream um and now we're trying to have this with this new system we're trying to introduce these same controls and we're under the same problems of once you already have this knickers bar what do i what do i do about that right like i have we had
old market that relied on the idea that i sold it to you at the value when you bought it that makes me whole and i don't have to think about you anymore right um and now we're trying to paste this thing on where we're like oh no hold on that that sticker's bar infinitely multiplies and therefore you never have to come back to me right um and that sort of substitute of product thing is at the core of copyright issues of the core bay i preferences um and we're very much in the process of figuring out what that even looks like in future what do you think the government should do whether in the UK or otherwise what types of policy positions are you have you not seen that you would like to see from political leaders i mean the first one's not really a policy change it's just not allowing for exceptions to existing like norms and laws like if a granny service and a scraper advertises that it's organizing human knowledge by bypassing paywalls and publisher blocks and then reselling that content onwards for its own profit then it should be shut down right like that's we have mechanisms
for that that's not a new crime right it's just an a new format instead we get 50 million dollars in funding on a series a from like well-known sources uh that pay for that activity and then we just look the other way right so like step one that's enforcing existing policy um copyright law has had issues for a long time that i'm not the most qualified person to get into right but i think there's a lot of very very smart people looking at copyright law and trying to figure out how to make it work for the future um and so that's a really important element uh making it very clearly illegal and there's some stealth crawler bills that are currently running um out here in the UK as well as New York state uh US federal government and a couple of jurisdictions are are looking at them that effectively just say you have to declare what your automated software is up to um and if you misrepresent that then that becomes a crime um i think it's an important if maybe a little performative uh because it's very hard to believe and very hard to enforce right but it's
important thing to set down as a precedent to say this activity is illegal right and this focus on it gets really critical like i've worked in ad tech for a long time ad fraud is now scraping by its exact same mechanisms um an automated thing shows up pretends to be something it's not does something that is against the terms of service or what should be happening makes its money and runs away right um we've looked the other way for a long time for ad fraud it's one of the big it's like organized crime revenue sources scraping is just that again and this is a time to properly take a look at that um and then lastly because it's going to be an important thing and it's what we're up to the same as done some good work on the so far we have to keep pushing on transparency um if these large companies are not going to be able to like sit down and actually have a conversation about a transparent liquid marketplace then that's what regulations for right and we've got to create the room for that so saying look it is a norm in a market not the seller knows when their
content is taken and it's what it's used what it's being used for um and there's real value in that right i cannot picture an environment where you cannot sit down as an AI agent builder and look at your agent be like man it would be so much better it had a little bit of this stuff you know and i could make so much more revenue if my agent was smarter at financial things cool let me call up john slate do a license deal give me some money the portion of my profit that my agent makes off it's using your stuff and we return to the market norms right like i think that's the place of a regulator is to start guiding us back to that norms so if they can then go back to being fairly hands off and let the market do it as well um but right now we just outsized influence of a bunch of capital it's gone into one side and it's a zero accountability when it comes to market norms um and then to me that's what government exists to correct um so last question looking forward what spurs timeline what goals are you trying to set let's just say the next you know couple monster year out yep in order to get
to the place where publishers brand feel more comfortable feel like they can start developing these marketplaces yeah so the standard itself is nearly out of draft form um so by the end of this month our v1 will be officially cut and out to the world but by the time this episode goes out you will probably be able to read it yeah fair shout yes or end of august i should be specific about that by the end of august it'll be a v1 um that's been uh we've had a good public comment period lots of engagement across the industry um that publishing will also have a nice thank you list in it which I'm as excited about as the actual standard um a lot of tech companies a lot of publishers a lot of brands have had a lot of say on making this thing good the next step is getting a spurt accreditation program together that actually it says and kind of rewards a lot of those tech companies that have worked on the standards with us for saying like look this is awesome you know you've done great work on this thing and now the rewards for that is that you build systems that are resilience and that we can then bring content owners to the table and say hey you can trust these
folks right like or more importantly you know what you're going to get when you work with these folks right you're going to get this kind of reporting back you're going to get this information um and starting to set those norms on the supply and the intermediary side uh we've got proof of concepts that are in the works that will show off um what this telemetry looks like with the reporting looks like uh you know and and contained more contained environments like it's not going to be chat to dt one day one sending out citation events um well we'll have working examples and that's partly going to be beneficial for publishers to find demand and markets to sell their content into uh it's also useful to show regulators and government that this is actually possible right and that there is a market over here and that if we had this information we could do some cool stuff um from a policy perspective right you can expect spread to start data a lot more active around um supporting publishers in the way that they go about understanding how to control and block how to understand bot access patterns and traffic even google which requires that if you want to show up on search that you have to allow uh them to scrape for their AI uh even go and as that starts to
change some right as as there's options now around which crawlers you allow and that sort of thing right like spural help with producing guidance on that and that sort of thing um four members and also also kind of more broadly uh and then like finally really you starting to think about how we use this kind of combined um not might not power but like the the world creates information and knowledge and content um and if we can be one part in bringing together this idea that that should be preserved and that we don't want to sacrifice that at the altar of like this thing so the words I could be using here um then this is a podcast as we can mark it as explicit if you if you if you'd like yeah I think finding um like finding the way to name and shame bad actors to elevate the ones that are doing the right thing um and to create a market that works for all of us is let's call that short and intermittent long roadmap right um I think getting more active particularly when it comes to things like like if you run a grounding service and a scraping service
at the moment I give you like a month to stop scraping stuff that you shouldn't be um because at the end of that we're going to start making a whole lot of noise um well that's per or that people like newscorp suing end users for using scraping services that are stealing shit like now you're you had your fun you've got your funding take your golden parachute right but we're going to find a way to make this market work um Alex bringer I wish you the best of luck thanks so much for chatting with me thank you thank you for listening to the media leader podcast our executive producer is me jack Benjamin special thanks to our production partner tri sonic you can find all our episodes on the podcast platform of your choice or on our website at uk.themediliter.com but do please remember to subscribe to be notified of our next episode you can also subscribe to our newsletter at the media leader to receive daily coverage of all the biggest news stories in media and advertising from all of us at the media leader I'm jack benchman see you next time
More episodes
More from The Media Leader Podcast

The future of media research and JICs - with the IPA's Dan Flynn and Graeme Grif...
The Media Leader Podcast

What is the future of freesheets? With Metro's Deborah Arthurs and Jo Mazenko
The Media Leader Podcast

The art and craft of audio advertising - AudioTrack at 10
The Media Leader Podcast

Is broadcasters' embrace of social video commercially sustainable? With Enders A...
The Media Leader Podcast