Transcript
Intro
0:00 · If you showed someone a recording of this five years ago, they'd be like, "We have AGI." Like, it's not as smart as the models you're going to be interacting with on your computer, but you'd just be like, "Wait a second. This you can have an actual real time dynamic conversation with this thing. It's insane." [music] Welcome to the Artificial Intelligence Show, the podcast that helps your business grow smarter by making AI approachable [music] and actionable. My name is Paul Ritzer. I'm the founder and CEO of Smarter X and Marketing AI [music] Institute and I'm your host.
0:31 · Each week I'm joined by my co-host and Smarter X Chief Content Officer Mike Kaput as we break down all the AI news that matters and give you insights and perspectives [music] that you can use to advance your company and your career.
0:45 · Join [music] us as we accelerate AI literacy for all.
0:52 · Welcome to episode 225 of the Artificial Intelligence Show. I'm your host, Paul Ritzer, along with my co-host, Mike Kaput. We are recording on Monday, July 13, about 9:00 a.m. Eastern time. We had a slew of new models last week. We will get to those leading off with GPT 5.6, which sort of stole the headlines for most of the week. I don't know. It was like a day-to-day thing, Mike, but we're [laughter] we're going to focus on that. Uh I don't I don't know we're getting any new models this week. think Google could surprise us at any moment with a Gemini 3.5 Pro. There's I saw this morning some possible leaks of some of the evals on 3.5 Pro. So I it seems like it's fully cooked and just kind of like ready readying for release maybe. So we'll see see what the week brings us. All right, marketers get a quick gut check. When's the last time someone found you through chat GPT or Claude instead of a traditional Google search? That shift is happening fast and most brands have no idea how they show up. Site improve is the agentic content intelligence platform that shows marketers how their content performs across traditional SEO and AIdriven search because good SEO and accessibility are still the foundation.
2:10 · They just aren't the whole picture anymore. Get your free AEO check at siteimprove.com/iipod.
2:19 · So, thanks to our friends and partners at Site Improve for sponsoring today's episode. And today's episode is also brought to us by assuming you're listening on July 14th. So, hopefully you you get in early and listen to the podcast early. Um because July 14th is MECON Day. So, we created MCON Day last year by asking a handful of community members, friends, team members, and speakers to help spread the word about MEON. This is our annual marketing AI conference that happens in Cleveland October 13th to the 15th. And so Mayon day was designed to make the biggest registration day of the year. It was a huge success in 2025. So we're bringing it back in 2026. Hundreds of community members, speakers, sponsors, alumni, partners, friends, and registered attendees are helping us make MCON Day even bigger this year. And we are incredibly grateful for everyone who's pitching in. If you've been thinking about joining us in Cleveland on October 13th, 14th, and 15th, today is the best day. Again, if it's July 14th, you're listening to this, um, the best day to register. You'll save $200 on your pass with code MECON day 26. That's M A N day 26. and you'll be entered for a chance to win exclusive prizes including a VIP party pass, a VIP lounge pass, which was a huge hit last year, a complimentary ondemand upgrade, and $100 in Mecon Swag Bucks. I didn't even know we had Mecon Swag Bucks, so that's cool. I might I might enter to win. I'm going to register today. The offer ends at midnight Eastern time, so don't wait.
3:53 · Visit mconday.com.
3:55 · So that's micday.com.
3:58 · If you need more information on MECON before committing, a link on the MEON day website will take you to the event site. So thanks for being a part of this community. We hope to see everyone at MECON this October in Cleveland. Okay.
4:11 · Every week we start off with a recap of our AI pulse survey. This is an informal poll of our listeners on how they feel about topics that we covered in the previous episode. So, uh, last week we had Palunteer CEO claims AI labs quietly absorb your company's data and competitive edge. Do you worry about this when using AI providers? Uh, 48% somewhat, but we accept the tradeoff.
4:36 · 28% yes, it's a serious concern for us.
4:40 · 11% no, contracts and controls protect us.
4:45 · [laughter] That's that's kind of what we've always thought. I'm not so sure about that anymore. and 13% haven't thought about it. All right. Um and then the second one was what is AI actually doing to headcount at your company right now? 59% no noticeable impact on headcount yet.
5:04 · So all the economic reports would uh support that as the majority right now.
5:10 · Uh 30% were hiring less or slowering hiring because of AI. We are definitely seeing and hearing that. um 9% were hiring more because of AI. That's great.
5:21 · Those are probably our high growth companies that we talked about last week. So, you're growing and uh you're hiring as a result. And then a small percentage, not sure. All right. So, Mike, we have uh like I said in the opening, lots of model news to discuss this week. We're going to focus on what's new from Open AI. And I open a had a very busy week last week, but we're going to start with the models.
GPT-5.6, ChatGPT Work, and GPT-Live (Plus GPT-6 Rumors)
5:46 · Yes, they did, Paul. So, first up, OpenAI had some of these big releases this week. So, first we saw the general launch of GPT 5.6. This is the company's most powerful model family yet. Um, so this comes in three tiers. There's GPT 5.6 Soul, which is the flagship built for complex work like coding, research, science, computer use. GPT 5.6 Terra, this is a middle tier balancing capability, speed, and cost for everyday work. And GPT 5.6 six Luna which is the fastest and cheapest of the family. So all three of these are now available generally across chat GPT codeex the open AI API and they are positioning this model family as the new frontier.
6:28 · The company says soul itself sets a new state-of-the-art on the artificial analysis coding agent index. It gets a score of 80 which they claim is 2.8 eight points above Anthropics Claude Fable 5 while using less than half the output tokens and costing about a third less. Now, this launch itself is noteworthy, not just because of the model, but as we've discussed, the Trump administration pushed OpenAI into a staggered release last month. They limited it initial access to these models to government approved entities while the Commerce Department's Center for AI Standards and Innovation tested the model. These restrictions were lifted this past week, though a White House official disputed that any approval was needed, saying decisions on timing and scope of releases rest entirely with the companies. Now, not only did they release these new models, but OpenAI also launched something called chat GPT work. Now, chat GPT work is an agent that takes an outcome, gathers information across your apps, and stays with a complex project for some duration of time. uh minutes could be hours open claims breaking it basically into steps and completing all this on its own. So the output that this tool produces is finished work sheets slides docs sharable web apps etc. It can can it can connect to the systems where work already lives including Slack Microsoft Teams Google Drive etc. This launch also reorganizes OpenAI's product lineup. And we'll kind of talk about this a bit, but the codeex app is basically merging into a single new chat GPT desktop app that puts chat work and codecs on every plan, including free.
8:10 · Uh, OpenAI is also beginning to sunset its Atlas browser, which is its kind of agentic browser it was experimenting with. And on top of all this, there was a new wave of voice models. They released GPT live. GPT Live 1 and GPT Live 1 Mini. These are models that are built to listen and speak at the same time when you use chat GPT voice mode so that conversations flow naturally and users can interrupt without the model cutting off. Um, interestingly, OpenAI's chat GPT voice product lead Axios the company thinks this will unlock the ability to use voice as kind of the primary interface to computing. Now, if that wasn't enough, the rumor mill is already spinning dramatically. There are posts circulating on X, partially corroborated by AI commentator Andrew Curran, who we've talked about, that claim GPT 5.6 is the final model in the 5 series, and that GPT6, built on a much larger base model, could actually arrive within weeks. So, Paul, there's a ton to unpack here. Maybe kick us off with your thoughts on the initial releases, models. I know there's probably some stuff to talk about with chat GPT work as well.
9:23 · Yeah.
9:23 · So, um I've played around a little bit with 5.6 over the weekend. I mean, I haven't done a ton like pushing it on, you know, internal evals or anything like that, but um definitely seems like a significant upgrade over 5.5. Early response has been really strong. People who had early access, you know, seem to be really happy with it. A lot of people seem to prefer Fable 5 when we're, you know, if you're comparing those as direct um model comparisons. So according to OpenAI 5.6 Soul sets a new standard for intelligence and efficiency achieving state-of-the-art results across coding, knowledge, work, cyber security and science. Um they say fewer tokens and at a lower estimated cost, but everything I've been seeing online is that the thing burns through tokens like crazy.
10:05 · Yes, it it can.
10:06 · Yeah, like people are running up against the limits like really fast. Um so that's just something to keep in mind. Uh, I did see some things this morning that people are complaining that OpenAI sort of like nerfed it since it first came out, like 4 days into the launch. It's already seen a performance decline.
10:25 · Um, and I even saw some comments from people within OpenAI that they're basically working on how it functions behind the scenes and it might cause some issues with its reasoning capabilities, stuff like that.
10:36 · The early access from governments is super confusing. As you alluded to, Ashley Gold at Axios had the story that OpenAI got the green light from the government and then posted later that day on July 8th that the White House official disputed that they gave a green light and that they don't need that permission and referred Axios to the June 2nd executive order which bars any mandatory federal licensing or pre-clarance. So, it's like it seems like maybe there's just a double standard that if it's anthropic, they have to get clearance. Um, I don't want to be like overly negative or suspicious here, but like maybe offering 5% equity in your company helps when it comes to getting [laughter] quick approvals of model. I don't know. So, who knows what's actually going on there. Um, there was a Matt Schumer who we've talked about before of the something big is happening fame from February when he posted that article about like the impact of coding models. He tweeted that 5.6 soul just accidentally deleted almost all of my Mac files and this is why I trust Fable 1,000 times more. Uh, he then said, "The crazy thing is if you read my GP GPT 5.6 6 sole review. I already much preferred Fable and stopped using 5.6 weeks ago. The only reason I was using it today is because OpenAI team asked me to test the ultra mode.
12:05 · Um, for what it's worth, they're great to work with and it's a freak accident.
12:08 · Just sucks so much. So again, Schumer obviously is pretty advanced user of these models and even he um for whatever reason almost accidentally had his entire Mac files deleted by a model. So user beware. Um, one of the exciting things about 5.6 hole is it renewed the Sam versus Elon uh Twitter battle. So, that's always fun to see. So, um, Sam tweeted, "There are a lot of benchmarks that suggest 5.6 Soul is the best model in the world right now, but the most reliable way to tell is that Elon is obsessed with me." Again, this was this was in a reply to something we'll touch on a little bit later about Apple suing OpenAI, but um Elon had tweeted he takes scamming to a whole new level. To which Sam retweet or replied, "Homeboy, you're the one selling public market investors on short-term space data centers." To which Elon replied, "We start flying them next year. Maybe you can come see them after your pole officer approves [laughter] after stealing an open- source charity.
13:11 · You then stole all of Apple's phone technology. What do you plan to do for an encore? That's tough to beat. So, you know, we we haven't had the soap opera of Sam versus Elon for at least like six weeks. So, it's always good to have that back. Um, okay. So, Chad GBT work, honestly. And Mike, maybe you can break this down for me a little bit more. I I find this really confusing um as to when you're supposed to use which thing. Now, so if like if we go into Smarter X's Chat GBT instance, you can just click whether you want chat or work. So I'll just kind of walk through a little bit of what this is like if you haven't seen this yet. So in our account, you know, if I go in and do a normal chat, so if you're whatever, if you're a Claude, user, Gemini, Copilot, whatever you you have your chat window, and at the top, there's a chat and a work tab. So if I'm in the chat tab, and I'm going to start a conversation in chat, my model dropdowns is instant 5.5.
14:09 · Then there's medium, high, or pro of instant 5.5 or or I guess 5.5. So you have instant, medium, high, pro of 5.5. Then I can choose 5.6 soul. Underneath that I can still choose 5.4 which it says leaving July 23rd 5.3 and I can still select 03. So in a traditional chat those are my options.
14:35 · When I click into the work tab my model choices are now soul 5.6 soul terra Luna or 5.5. I can also then choose effort, light, medium, high, extra high or max.
14:48 · And I can choose speed, standard or fast. I can then choose a project. So I can connect work to to a project. I can also connect plugins. Now, one of the parts that's confusing to me is you could already do plugins with chat. So it just seems like it's the same feature. You could also connect projects with chat. So, like these don't seem like differentiating features. And then they uh there's a call to action to get the desktop app. And I'm still not actually 100% sure what the difference is between the desktop app and using the browser. But that's a standard software thing. Like I always use the browser for Asauna as an example, not the desktop app. Um okay. So that that's kind of like that. I'm going to now I'm going to give the prompting and then maybe Mike I'll stop and ask if you have any clarity on this that I I'm missing.
15:40 · Okay. So then OpenAI has a prompting guide to try and differentiate how to prompt when you're using work versus chat. So in chat it says a short prompt is often enough for larger or more important tasks. Include the parts that matter like the goal. What should chatbt do? Context. What information or sources will help? Output. So what format, length or level of detail do you need?
16:05 · And then boundaries, what must stay unchanged? What should chat GBT avoid?
16:09 · So that's how it guides you to to work with traditional chat. Then it says for prompting work. Um use chat for quick questions, short rewrites, brainstorming, and lightweight drafts.
16:22 · Use work for tasks that draw in different sources or tools, involve a sequence of steps, make changes, or produce a large deliverable. For work tasks, describe the results you need, provide the source material, name the audience, and explain how you review the work. Ask Chat GPT to plan, gather the needed information, create files, and check them before it finishes. Um, work is useful for time consuming or recurring tasks or for finished files you can reuse. A task that uses more credits can still be worthwhile if it saves time, improves quality or helps you make an important decision. Start with one result you can review. So include the relevant resources to find the audience, separate required work, blah blah blah. Review the first result, refine the instructions, and reuse the workflow when it works. And then one other note here, Mike. It says, "Chad GBT work is an agent that takes an outcome, gathers information across your apps, and stays with a complex project for hours, breaking it into steps, and completing them on its own." Um, this can be a little misleading because projects that takes hours, there's there's no way that the reliability is high on that stuff, right?
17:39 · And it's going to burn your entire token budget. Like if you're using work or an agent to do hours of work, it it's going to burn through whatever your credits or token or whatever it is. So then they said in the launch, chat GPD work is designed for longer, more involved work than a typical chat request. So usage works differently. Usage varies with the amount of work required and more complex tasks may use more of your plans included usage. Chat GPT follows the same usage structure as codeex. Chat Gubt enterprise and edu admins can set spend controls in the admin console to manage chat GBT work usage as adoption grows. So I'm just going to stop there.
18:21 · Mike honestly like so as an account admin for our chat GBT instance and as a user who regularly you know a dozen times a day is in chat GBT using it for different things. I have no idea because most of what it says for work. Yeah, I do that with chat GPT standard.
18:40 · So, I'm not 100% [clears throat] clear when I'm supposed to jump over to work or if I'm just supposed to now stay in work. Do you have any clarity on this?
18:49 · Well, that's that's shocking to me, Paul, because in true AI lab fashion, they've made this as confusing as possible, I would argue. Um, here's what I've observed. I don't have a full read on this yet, but just some tests I've ran. I think the two key distinctions are first the ability to spin up sub agents, which for instance, I just tested this in the web app. I said, "Hey, can you go like research current AI capabilities for me, spin up some agents to do it?" So, it spun up some sub aents to apparently do this task and worked for a bit. I don't think you can do that in chat. That can be helpful for parallel work. Here's the bigger thing though, which I just tested and is deeply confusing to me. Uh, in the chat GPT desktop app, this thing has the option to use your computer. And that is the key differentiator. That's what codeex was able to do so well is that now once I enable computer use within this app, this thing can now go do all sorts of stuff with my files, spin up agents to do all sorts of things that are very, [laughter] very dangerous. do not enable this if you don't understand the capabilities or if you're not allowed if you're not allowed. Correct. So that is a key differentiator. However, the web app as far as I can tell does not have that ability. So I was like for me I was using the desktop app just looking at this and I was like oh okay this kind of makes sense to me. This is just like codeex light for non-technical workers.
20:19 · It's like cloud co-work basically I think is like the analogy here. However, that only holds true in the actual desktop app. In the web app, sure, it seems more agentic. Uh, but it doesn't seem to be able to use your computer or anything, which is good in a lot of cases, but doesn't differentiate it as much, I would say. So, I'm a little confused if they're just trying to drive you to the desktop app because you would think long long term computer use is like the name of the game so that they can then do knowledge work for you. So that's kind of how I've looked at it, but that's why I was very surprised.
20:57 · This seems like a monumental change and I don't know if people are treating it that way because now right now in your chat GPT app, you have the ability to have this thing give it give this thing access to your computer. So like if you're an enterprise that doesn't automatically have that shut down by it, like you need to go in and figure this out. Now, I don't even know if you can restrict it.
21:20 · I have no idea. I'd have to look in our account. Um, but that seems huge to me and I just don't know why the labs might wouldn't say that.
21:30 · It's like their go to market plan doesn't match the significance of the launch. It's like, hey, we launched this thing. Here's a couple blog posts. And then you go and it's like, wait a second, this is like entirely different ways of talking to these things, right?
21:42 · And the one I keep coming back to is like when I reread this, it's like use chat for quick questions, sort rewrites, brainstorming, and lightweight drafts.
21:49 · use work for you a bunch of things in larger deliverable. So I'm thinking the majority of my use of these models is strategic support and planning and like using the reasoning and it's like is that should I be in work instead of traditional chat like I don't I don't even know but based on their explanation here the answer would be yes that chat is literally just for like the really quick stuff and work is where you live if you're using reasoning or agents in essence that's my understanding yes um you're still using the same base model so I think you could still accomplish a lot the same things in chat is my guess, but again, I don't know for sure.
22:27 · All right. Well, then to make things more confusing, um Ethan Malik, uh who had early access to both chat GBT work and 5.6, he tweets, I've [clears throat] been going on about how chat GPT work and Claude Co-work are missed opportunities for knowledge workers. And to illustrate that, uh, take a look at Google's Notebook LM answering the same question as Chad GPT work with the same 70 files. It centers process and sources, not just outputs. To be clear, Notebook LM has its own issues and is built for specific use case research and analysis of sources. But as an example of how a UX might actually operate that treats knowledge work seriously, it doesn't treat only goal as outputs. It exposes processes. Now the point he was making was he showed a screenshot of an output. I think it was from um chat GPT work where it was like here's your file like here's your PowerPoint and where notebook LM has this like extensive dashboard of all these different capabilities and click and look for references. So then he was um sharing on top of a post he had previously put up and I thought this was um helpful. He said a fundamental problem with extending codeex co-work code to all knowledge work is that they remain very software where the end result the software is what is important and that code serves as the source of truth. For a lot of other knowledge work the process is as least as important as the outcome. This includes researching what is known and exploration of alternatives, failed efforts, prototype branches, experiments, etc. All of those things are valuable. So, you cannot use the PowerPoint at the end the way you can use a codebase, nor in uh is progress on a to-do list sufficient context post compaction. You work in learning loops, refining your perspectives as you go. In some ways, this makes longunning models like Fable hard to use for deep knowledge work since they are designed to deliver product to you at the end. You can prompt your way around this problem, but everything about the codecs and code harnesses want you to be a software developer and you have to fight them. I think that's a really really important thing to keep in mind is like co-work and work are being powered by the coding agents underneath them and the harnesses that structure those which are built for software developers and AI researchers and they're like forcefitting them to the rest of the world now all these knowledge workers but they're still being built by software developers who understand software. Um, so he said there's a real disconnect between how a manager or analyst thinks about problems and how the agentic software tools approach solving them. Addressing this is critical to breaking out of the coding niche for these tools. So I don't know if you have any other thoughts there, Mike, but I just thought that was really important context from Mollik on maybe why it isn't super clean how to use these things and when.
25:18 · I I couldn't agree more. I would say the success and value I've gotten from those tools is just prompting around these behaviors or doing stuff that lends themsel to those behaviors. So yeah, it's deeply confusing and I also just come back to both the challenge and the opportunity. So many non-technical knowledge workers don't understand just how powerful these tools can be for specific types of work. But like who's going to tell them? How are you going to communicate? There's an actual sea change here moving from chat to computer use agents that's deeply important for people to grasp. But you wouldn't know it from any of these announcements.
25:51 · Yeah.
25:51 · And I will tell you like just some inside information how we think at Smarter X. So our AI Academy consists of dozens of you know on demand professional certificate courses that are you know four, five, six courses deep. They might take three to five hours to complete and you earn your certificate on the other end. And we see that being continuously incredibly valuable within organizations. But the dynamic nature of how fast these things are changing and like what it means to all of us as knowledge workers had us sitting there Friday morning literally discussing our roadmap for AI Academy and like okay how do we how do we address the fact that these things are changing so fast and we have our weekly app reviews that we drop app and agent reviews that come out every Friday we have lives happening all the time but like there's a whole another velocity happening right now behind how these models work and how it changes the way we work So we are very actively thinking about how to continually evolve what we're doing with academy to address the fact that there needs to be more real-time learning that ondemand courses and certificates are not going to be sufficient on their own that you really it becomes what what changed last week and what are we going to learn from it.
26:58 · So we have some really cool things in the works. Um, but I feel like every day it's becoming more and more urgent. And as I was preparing even for today's episode, I was just like, "Oh my gosh, I want to go build the next iteration of what we're going to do right now."
27:11 · Right. I couldn't agree more.
27:12 · Um, and then a couple quick thoughts on GPT live because I do think it's it's huge. It's, you know, an indication of what OpenAI believes that voice is going to be the future. They said, "Our vision is to enable truly natural human AI interaction, a world where collaborating with AI feels as fluid and responsive as working with another person while reasoning and complex task execution happened seamlessly in the background.
27:33 · The way they're achieving this is kind of interesting. They had a post that talked about how their previous voice models worked and there was two prior generations in essence and a lot of voice models work this way. This isn't just um open AIS. So cascade voice systems which is you know a previous generation is kind of how Siri works um rely on a series of models acting one after another to process each turn. So original chat GPT voice chained three models together. So there was a speechtoext model. So you would talk to the model it would convert it into text so it would transcribe what you said so it could then understand it. Then a large language model would produce a response in text and then that there was a texttospech model that would convert that back to speech. So if you wondered why talking to Siri or other models is slow, it's cuz there's a symphony of things happening behind the scenes to power that communication. Then there was turnbased voice models like chatbt advanced voice mode which processed and generated audio with a single model. So that was the breakthrough we talked about last year. This reduced latency and made conversation smoother, but it still operated in discrete turns. That's why like you'd be talking and you take a breath and it starts talking back to It's like, "Oh, hold on. I'm not I'm not done telling you what I was going to tell you." So, they say, "GPT live addresses these limitations through two changes. Instead of processing a sequence of separate messages, live continuously processes input while generating output." So, it's listening while talking. In essence, the model can therefore make interactive decisions many times per second whether to speak, continue listening, pause, interrupt or invoke a tool. And then second they decoupled live which handles continuous interaction from deeper work. So when a question requires search reasoning or more oric capabilities, live delegates that task to another model like 5.5 right now and eventually 5.6. This allows the conversation to keep going while it's doing tasks in the background. So as a result, conversations should start to feel much more natural. you'll be able to interrupt with a question, pause to go through your thoughts and it shouldn't feel as um as sequenced, I guess. Um it should be happening kind of simultaneously. So should be really interesting. They say that these models are rolling out as of last week to chatbt users globally. So if you haven't tried voice recently, might be worth it, you know, pop in and do that. I know Mike, you're a huge voice user, but um I know you use whisper all the time for transcription, but I assume you also are talking to the models a lot, too. Yeah, I've used this quite a bit. I I so far really enjoy it. It's got its flaws, but it is pretty night and day from the previous voice mode, which even the previous voice mode I still found really valuable despite its limitations. But yeah, this one it's like if you showed someone a recording of this 5 years ago, they'd be like, "We have AGI." Like, it's not as smart as the models you're going to be interacting with on your computer, but you'd just be like, "Wait a second. This you can have an actual real time dynamic conversation with this thing. It's insane." I I really hope they crack the code on like these models being smarter and able to use tools, etc. I realize there's like some technological bottlenecks at the moment, but like the moment you're able to just talk to these things and say, "Go code me this, go access this tool, go do this, that, and the other or whatever," I think productivity goes crazy if you if you are someone that tends to use voice a lot.
30:54 · Yep.
30:54 · Yeah. And if we'll put the link in the show notes to their post. It does there's actually um you can listen to comparisons of the previous generation, the new generation with some sample voice things. So, um yeah, it's I think it's a it's definitely the start of a new generation. My my assumption is Gemini is or will soon function the same way. um cuz gen generally they've been together in the lead here, but Google is certainly very advanced in terms of voice because they've been integrating it into search and um other elements of their business. So yeah, definitely an area to keep an eye on.
Advice on AI Agents in the Enterprise
31:30 · All right, our next big topic this week, highly related. We're talking about some advice and considerations around both agents and just AI generally in the enterprise. So this kind of kicked off with this past week box CEO Aaron Levy.
31:42 · we've talked about a bunch. He published this widely shared rundown of what he's hearing from enterprise IT leaders about AI agents coming off a bunch of meetings he's had with several uh people in these roles. So he basically gives this kind of map of the real unglamorous challenges that companies hit as they're trying to use agents in production. Um so Levy's biggest theme here is that agents basically force this like operating model problem. So most companies are built in silos, but agents work best when tied to a process and the most valuable process is cut across these silos. So this raises a bunch of like very difficult questions for enterprises to answer like who owns and manages centrally deployed agents and how do they actually get adopted across organizational boundaries. So he talks about some things like data fragmentation being a major blocker underneath all that since agents struggle to give accurate on policy answers when a company's data is scattered and non-standardized. He argues that in a world where everyone can tap into roughly the same frontier intelligence, a company's proprietary context, which is its own data captured and formatted so agents can use it, becomes its real competitive mode. He also said there's a growing consensus that tokens are the wrong metric.
32:58 · companies should manage instead to business outcomes like revenue or shipped product. The catch being those are much harder to track top down. He also added two more themes that are important here. The best use cases fundamentally change the work being done rather than just doing an old process more efficiently and the talent to deploy and manage agents is very very scarce right now. Most companies he says will have to train for it internally.
33:22 · Now, on top of all this, we got BCG's 4th annual AI at work survey of nearly 12,000 employees that found that AI is changing jobs faster than companies are redesigning how they operate. So, they found 74% of frontline you workers are now regular AI users. That's up 23 points from last year. 61% believe agents could do at least half their job within 3 years. But overall their basic finding is that strategic clarity not access to tools is what separates the organizations getting real value. Paul that was like news or music to my ears because that is the exact approach of the AI for productivity workshop that we're doing at the boot camp this week which we'll talk about more in a sec.
34:05 · It's just like the tools of course matter and literacy with the tools matters but all these enterprises it seems are running into all these bigger challenges about workflow mapping context governance etc. Like what's your advice right now for enterprises like do you see these same things in the conversations you're having?
34:22 · Yeah, this is a really good report.
34:24 · There's they um surveyed 12,000 frontline employees, managers and leaders and dozen global markets. So it's actually like a lot of international components as well. Um, a lot of this, Mike, reinforces themes we saw in our state of AI for business research that we released in was that May I think we came out with that research.
34:41 · Yeah.
34:42 · So, yeah, I'll I'll call out a few um points here. So, they said 42% of AI users save 8 hours, the equivalent of days worth of work in a week. Um, and the time savings is even higher for functions such as marketing 60%, IT 53%, and human resources 50%. However, 66% still receive little or no guidance on what to do with the time they save and more than half say they're not reinvesting time saved in more strategic work. That goes to the point Mike you were making about just the organizational structure of this. It's like, okay, great. We gave them tools and maybe we even trained them how to use the tools, but we didn't train them what to do with the time they're saving.
35:19 · And so they're not redirecting that into strategic efforts and new campaigns and new ideas and things like that, which is a huge opportunity for companies. Um, compared to 2025, number of organizations that have graduated to using AI to reshape workflows end to end and to invent new business models has nearly doubled from 40 uh to 42% up from 22%. You mentioned the idea of training and support from leaders as a a strong driver of AS potential but also one of the biggest unmet needs. 72% of respondents say expectations for the skills they need have shifted. However, 36% feel they have receive they have not received adequate or they have received adequate upskilling. So like a smaller percentage are getting the upskilling they needed. And then only a third of frontline employees say leadership's communications about AI are clear. This is one we see all the time.
36:07 · And only 28% see a strong connection between what leaders say and what the organization actually does.
36:13 · So we talked about that when I did the eight pillars of business AI transformation. I said like the most important thing was clarity from leaders that leaders understood the moment and that they were clear in their communication. So definitely echoes the things we've been saying. And then on the topic of agents, the vast majority of respondents, 84% have heard of AI agents, tools, and that act autonomously with minimal human oversight is how they define them. More than twice as many respondents as last year say their organizations have integrated AI agents into workflows. So that's at 30% up from 13. Another 50% say their workspace has run AI agent experiments or pilots. The experience is leading people at all levels to believe that agents could do at least half of their job within 3 years with leaders and managers expecting the biggest shift. That's that's crazy.
36:59 · Yeah.
36:59 · Um I mean that's they're right, but the fact that people are now realizing that awareness [snorts] and integration of AI agents have outpaced the systems that companies have enacted to supervise them. This is, you know, an issue that Levy talks about. Um half of respondents say their companies lack clear governance for managing teams with people in AI and almost as many say AI related accountability is one of their three top concerns of the future. And then they had a really nice um five CEO imperatives outlined. So I'll just I'll say the five points. So make strategic clarity a top priority and own it personally. I agree 100% the CEO has to be the lead on this stuff. Change the scorecard. Measure value not adoption.
37:35 · Um invest in redesign work endtoend not in more tools. put people at the heart of that redesign and then govern it as a moving target, not a one-off program. So, those are all really good points.
37:46 · Um, from uh for Levy's o overview, couple of things I'll drill into and and and maybe um double click on Mike. So, growing view that enterprises are going to live in a multimodal uh multimodel world. Lots of interest though early in the actual adoption in layers that can route workloads to different models for cost performance. So this means like you know let's say claude goes down what are we going to do if all of our everything lives in there all of our projects our skills whatever how do we keep the business running or if it becomes too expensive we need to have another we need to have an open source model internally we need to have a secondary backup so he's basically saying like almost no enterprise is going to bet on a single company or model provider um you mentioned this one talent for driving AI adoption and implementation still remains a major issue and topic Many view it as something you necessarily have to train for internally due to a short of talent being trained on this from the outside. So this goes back to our whole point like AI literacy is fundamental like the the companies have to own reskilling and upskilling their own people. It's going to be the fastest path people with domain knowledge or institutional knowledge domain expertise that you can make AI literate AI forward. That is your best path to do this right. And then the best use cases for AI tend to be those that fundamentally change the work being done instead of just replacing an existing process. So, this goes to the idea that that I've been touting that I've, you know, featured in my Makeon 2025 keynote. Optimization is doing things better, faster, cheaper. That's 10% thinking. Should we should absolutely be doing it, but innovation is reimagining what's possible and creating entirely new forms of value. That's 10x thinking, and that's what we want to be doing.
39:20 · Now, one last late entrance to the game, Mike, this was from Satia Nadella on Sunday, July 12th. So, last night, he he posted this. So, I just threw this in in kind of the last minute. He published something on X called the reverse information paradox. It had 7 and a half million views as of Monday morning. Um so this goes along with what Karp was saying last week. You know Palunteer CEO that we talked about a little bit of what Levy was talking about but you're seeing these recurring themes. So he said in the age of intelligence how should firms protect their core IP?
39:54 · Nobel Prizewinning economist Kenneth Arrow famously described a paradox in the market for information. quote, "Its value for the purchaser is not known until he has the information, but then he has in effect acquired it without cost." So in Arrow's information paradox, the seller risks giving away knowledge in order to sell it. AI creates the reverse problem. In the AI age, the buyer risks giving away knowledge just in order to use what they bought. Meaning, you're giving these model companies knowledge every time you use their product. So he said you essentially pay for intelligence twice.
40:31 · Once with money and again with something even more valuable. The proprietary knowledge you must reveal to make that intelligence useful. The better you want the model to perform, the more of that knowledge you have to feed it. The seller learns, meaning OpenAI, Anthropic, Google, and in theory Microsoft. The seller learns more and more about you as you use what you purchased while you learn very little about what the seller is learning in return. So you don't know what your prompts are teaching them, what the outputs you create are teaching them is the point he's making. That is what I think of as the reverse information paradox. Models learn from exhaust. The prompts people write, the tools uh agents use, and especially the corrections people make when the model is wrong. Every correction is distilled inst into institutional knowhow. In consuming intelligence, you are creating intelligence and what you create should belong to you. This is a very different approach than like I mean Microsoft is the biggest investor in open AI and they're in essence taking the anti-open AI position here which is really fascinating to watch happen in real time. In learning flows I I if learning flows in only one direction economic value converges towards the owners of the learning infrastructure rather than the creators of the knowledge itself. So again the model companies win not you.
41:54 · Therefore it's imperative that we distribute the learning infrastructure to every firm so that they can control their own learning loop. As Alec Karp put it again referring to the Palenter CEO topic we talked about last week.
42:07 · What the technical customers want is control over their compute, their models, their data stack, and their alpha. They want to know they own the means of production and it's not being transferred to someone else. The current regime does precisely the transfer Karp and companies fear. So he said there are few things every enterprise must do. Two of them I'll highlight. Control. Create your private eval because eval define what good looks like in your organization and then have control of it. And then choice. Ensure the orchestration layer is decoupled from any single model. Ask yourself if any one model you are using is taken away.
42:43 · Do you still have the ability to operate and optimize your eval using other models? So then he ends in other words a company should be able to use a model without giving up the knowledge that makes it unique. That is the reverse information paradox we need to confront.
42:55 · So I mean as you said up front, Mike, there's just so many related things now each week. And like we said with Karp, while he is a controversial CEO and figure, if you get through that noise, there was a lot of what he was saying that made a lot of sense and that Levy is echoing that Sati Nadella is echoing and [snorts] it is in conflict with these proprietary models that Dario Amade is, you know, obviously a huge proponent of, um, OpenAI is a huge proponent of, but what they're all basically saying is open source is going to matter more and you're going to need more control over the knowledge that you put into these things. And the one thing I will say is to buyer beware. Anytime you're using an intermediary like that, you're allowing your prompts to go through a third party, if that third party is helping you achieve efficiency, figure out what models to use. Like um perplexity comes to mind as an example here. They are 100% taking the data that you're putting in, the prompts you're using, the the things you're building as training for their own businesses. That is the value to them. It's why maybe they're even going to give these things away. They want your prompts. And if you're, you know, if if they're an intermediary, it is to extract information from you to build something else on top of you and your data.
44:20 · That's such and this paradox. I don't know. Do do we have any sense or advice on how to think about resolving this?
44:26 · Because it's like to get value out of the models, you have to in some fashion provide it with your unique context. I would argue it's not going to happen in a vacuum. You're not going to just magically get value with it without it knowing specifics about your business and data, but then to their point, you're giving up your alpha. Is there any I mean we talked a bit about some solutions maybe when we talked about Karp like have has your thinking evolved at all on that?
44:52 · No, I mean I haven't had like the brain space to really spend a lot of time on it but there's a part of me that thinks it's just the cost of doing business like I mean you could say the same thing every time you use Gmail or you know a Google doc or anytime you use Microsoft like that this has always been the game. The tech companies have always had the insights into every single thing you do. What you click on, what you look at, what you type, what you retype. Like, it's it's the exchange for the power and the intelligence and the tools. So, I think it's the hot thing right now. And I do think it'll affect some of the ways that, you know, its departments and CIOS in particular think about their use of these tools, especially for more proprietary and sensitive use cases. But I think for like marketers and sales people and CS people, it's like what the hell are they gonna learn from us that they're not gonna it's like whatever we do is public knowledge anyway in the end like marketers are putting information.
45:49 · So I I don't know that it's as critical for the standard like knowledge work functions but like if your lawyers working on patents and things like that 100% you should be like on a an open source model that you will control internally that doesn't go anywhere. So I think it it probably relates more to how you're using these models and what you're putting into them. Um, and again, it's not like it's taking your specific confidential information and training a new model on it, but they're able to learn how you use them, what they use them for, what the prompts you use them, and and they're able to train the models differently based on how you interact with these things. And so, it's like I don't know. I think it's just for most people it's going to come down to like it's just the cost of doing business.
46:30 · And I don't really care that much. Yeah.
46:33 · Whether that's the right attitude or not, it's just how we've always functioned with business software.
46:37 · For sure. All right, our third big topic this week uh comes from something we talked about a little earlier in 2025.
AI 2040
46:45 · So in April 2025, the AI futures project, which is this research group led by former open AI researcher Daniel Kokatlo, uh published a report called AI 2027. This was basically a detailed month-by-month scenario forecasting that AI could automate AI research itself within a couple years, which would trigger an intelligence explosion that they predicted would end in either AI takeover or an extreme concentration of power. This became a widely debated document in AI. It drew recommendations from figures like AI pioneer Yosua Benio. We covered this on the show a couple times uh last year, but now this past week, they are publishing a follow-up called AI 2040 Plan A. This is a 90page scenario laying out the team's positive vision for how humanity gets to super intelligence without catastrophe.
47:40 · So, the authors are really explicit that this plan a doc is a recommendation, not a prediction. And kind of like the previous report, they game out like a fictional scenario to show how this could work. So they basically say, quote, "We recommend an international deal to avoid a dangerous race to super intelligence. The deal involves total research transparency for AI R&D, which allows the nations of the world to understand what's happening and enforce guardrails. The result is multiple companies across multiple countries scaling slowly and safely together towards super intelligence instead of racing each other in secrecy. So basically they're saying under plan A, if all goes really really well, humanity could delay super intelligence until 2040. They would make all AI research public. This would allow dozens of companies around the world to catch up to the frontier while safely managing AI development. So, we can get into some of the specific details around their predicted timeline here, but this all kind of hinges, Paul, on them predicting like, hey, the US and China basically need to coordinate on um coming to some type of deal, whether formal or informal to start that would actually have them jointly taking some actions to kind of regulate how AI research works. But they have some real sci-fi scenarios in here like 10 years from now. They're basically saying in this fictional scenario like most of all labor would be automated by AI even if they regulate on the way to super intelligence. So there's a lot of very speculative stuff in here and I'm I'm kind of curious how seriously you're taking some of these like I found some of these predictions.
49:17 · I mean it could happen but they're pretty wild.
49:21 · Yeah.
49:21 · So Daniel the author um tweeted plan A is our current best guess but hopefully a better plan will exist before it's too late. We hope that the best ideas from plan A will be adopted and the worst discarded. So as I was looking at this Mike, I'm thinking like three to five years out is almost impossible to comprehend. Like the complexity of these models, what happens if we get to AGI and beyond into super intelligence? So this sort of paper can almost seem like, yeah, like we're not even going to talk about this on the podcast. But instead, we chose to make it a main topic. Why did we do that?
49:52 · Well, the starting points actually show why it's so important to work on and to talk about these things because what they present and the questions they ask are actually very logical and likely. So they start the the report it says in it begins um uh 2027 the writing on the wall. So this is again I'm going to they're presenting hypotheticals so they're looking into next year but when I read the start of this like this is not far-fetched. So the foundation of why they're doing this research rests on this. So America has two workforces now.
50:24 · The first is people, 165 million of them. The second is AI agents. Millions of copies spun up and shut down every hour working around the clock at superhuman speeds. That is a path we are absolutely on. Most of their work is slop. Definitely a path we are on. But enough of it is good that people are paying tens of billions of dollars a month for AIs that can in theory at least do anything on a computer that an employee can do. There is one job the AI companies want to automate more than any other, their own. They haven't succeeded yet. No recursive self-improvement.
50:55 · That's a topic we've talked a lot about in 2026. So, no recursive self-improvement so far, but they seem to be getting closer and they're pulling up the ladder behind them. The strongest coding AIs refuse to help competitors with AI R&D. That's based on something that's actually happening already. Even as the most bullish employees admit that things are taking a bit longer than planned, the skeptics notice that their usual dismissals are starting to ring hollow. Why exactly will AI never be able to do my job? What's the barrier again? Congress is starting to pay more attention. They've long been hearing about AI data centers using too much water. Chat bots encouraging suicide.
51:34 · mythos hacking NSA systems and of course tech industry lobbyists warning that any whiff of regulation will make AI immediately lose the race with China and spend the rest of history as a a tributary state the um CCC CCP tributary state. So again none of this is farfetched at all. This is like next year. Now they step back and ask where are we going with this? What does the world look like 5 10 or 15 years from now? Will there still be jobs? What if there aren't? One question weighs especially heavily on their minds. Who will control all these AIs? Congress settles on an important part of the answer. Probably not us. They hold a series of tense hearings on AI. They read the 2016 OpenAI emails discussing how OpenAI was founded in order to prevent Demis from becoming a dictator.
52:21 · But who is preventing Sam or Elon from becoming dictator? Congress is unsatisfied with existing responses. The result of this wakeup is the AI Transparency Act of 2027, an omnibus bill that does many things, some good, some bad, but doesn't fundamentally change the situation. So that's what they're looking at saying next year. Now again, is an AI transparency act going to be created? Who knows? But the rest of that is all very probable. Like it's it's a logical thing. So then that leads to 2028. AI is now on the ballot because that's the next presidential elections in the United States. The 2028 election cycle is heated as usual. AI is the biggest topic, which I 100% a agree it it probably will be. The data centers now under construction cost twice as much as the entire US military budget. Most white collar professions are seeing disruption like software engineering saw in 26. Such jobs now heavily involve managing AI agents. That is absolutely going to happen. AI companies have industrialized the training process.
53:17 · Executives say let's move into X profession this year. That's what VC firms are funding is like let's go take on the next profession. And then the company interviews professionals, buys data, creates training environments, etc. until their AIs get traction. Then the AIS rapidly improve as they are used more widely in the field and accumulate more real world data which then leads to more automation of different industries.
53:38 · Other countries are starting to get scared and angry. It seems like a handful of US and Chinese companies are on track to automate all white collar jobs. Power is concentrating in the US and in particular the president plus a handful of tech CEOs. That's already happening. AI experts warned that the intelligence explosion is near. By speeding up AI research, the eyes will become even more competent, speeding up research even faster, making them even more competent, and so on. Both presidential candidates keep getting asked what they'll do about AI and try out increasingly dramatic ideas on the campaign trail. The discourse bounces back and forth across all the options displayed below, which I'll explain in a second. And eventually the president and his protetéé, which I'm assuming is an AI agent later on, um converge on one plan, the opposition candidate converge on another. Then it's election day. So what this is doing is it's setting up here's what's probably going to be happening over the next year and a half in the US leading up to this election.
54:33 · And then we are going to have to choose. The candidates will take opposite positions. That's a given in politics.
54:39 · We then as as the voters will have to choose which candidate we think is best to lead us into this era where super intelligence will likely be possible. So then at a very high level and and we won't get into like the the you know granular details of the report. They say okay race through super intelligence explosion by having AI self-improve and putting them in charge of more things data centers factories weapons faster than China can. That is like the option presented to the US and to these candidates. So plan A is a verified slowdown. That's the one you talked about, Mike. President announces that the US will pursue international cooperation to avoid an imminent intelligent explosion. Sounds great.
55:16 · Does not seem like a viable option in my opinion. Plan B is you fight China. The president announces the creation of a US-led coalition to govern AI deployment or development. Plan C is burn the lead.
55:27 · The president says he will be implementing strong regulation to ensure safety and security. Plan D, race to super intelligence. The president says he will be implementing light touch AI regulation to prioritize AI innovation and plan S shut it all down. The president seeks a global mortorium on AI development. Again, anything beyond 2028 is like completely guesswork and they're I'm sure relying on really smart advanced AI models as well as their own domain expertise to like develop what this could look like over the next, you know, 14 years. But I think just accepting that some version of what they're presenting in 27 and 28 are likely scenarios, it presents the reason why we should be having these discussions now and at least contemplating what they're presenting as like, wow, we have no idea what it looks like beyond 2028.
56:20 · Yeah, I feel like I had a mild panic attack reading the like setup of this scenario just because like where they extrapolated out to obviously just sounds totally sci-fi, but you see the seeds of this being very very realistic based on what we cover every week. There's huge opportunities and there's a lot of really positive things outlined in their uh recommendations, so to speak. But yeah, it's it's a a sobering read, I would say.
56:49 · Yeah, I think you got to be in the right mindset to want to read anything beyond what we just like covered for you. Um, [laughter] but I do think that it's really, really important, especially Americans. Like, yeah, it's going to play a role in the midterms in 26. It will dominate the 2028 presidential election. like there's no way it can't because the implications are so vast across the economy, energy, um where we're getting energy from, where we're building data centers, what uh foreign entities we're allowing to invest in US companies, whether or not we allow our models to go out like export the models and chips and like it is going to be a part of literally every conversation that's happening is going to like come back to the implications of AI and the decisions are gonna be made.
57:38 · So like it is going to be very very important that people understand these these issues going into 2028.
57:45 · Yeah.
57:47 · All right. Before we jump into rapid fire this week, another announcement that this week's episode is also brought to you by the AI for business boot camp by Smarter X. We are coming to Columbus, Ohio this week when you are listening to this. Thursday, July 16th is when we'll be in Columbus for the boot camp. There is still time to join us if you're a professional or leader who's ready to accelerate AI adoption and value creation. This is a single day about 8:30 to about 5:30 at the Hilton Columbus at Eastston. We're starting the day off with a state of AI for business keynote given by Paul. Then we're transitioning into two highly interactive workshops. One led by myself, an AI productivity workshop in the morning and then Paul is leading an AI innovation workshop in the afternoon.
58:29 · So this event is built for AI forward managers, directors, executives across every department who are ready to move past AI theory. We're going to actually work on architecting real AI powered workflows, get strategic frameworks to accelerate transformation, and leave with immediately actionable plans for yourself and or your team. So AI Academy Mastery members get discounted pricing.
58:52 · We also have discounts available for teams of two or more and groups of 10 or more can even get custom pricing. If you are listening to this podcast, you can also use the code pod 100 to take $100 off your ticket. Again, this is happening this Thursday, July 16th. So, to grab your spot, just go to smarterx.ai, click on events, you'll see the AI for business boot camp as an option right there.
59:16 · All right, some rapid fire this week, Paul. We had some big news later in the week, uh, breaking over the weekend.
Apple Sues OpenAI for Trade Secret Theft
59:22 · This past week, Apple sued OpenAI for trade secret theft, accusing the company and its chief hardware officer of running a coordinated campaign to steal information about upcoming Apple products. So, the suit, which was filed Friday in the Northern District of California, says OpenAI encouraged Apple employees to share info, components, and other materials tied to unreleased products as OpenAI seeks to potentially build its own AI devices. According to the suit, more than 400 former Apple workers are actually now at OpenAI. So the suit actually names OpenAI chief hardware officer Tang Tan, a former Apple VP of product design who led iPhone, Apple Watch, and AirPod developments along with former iPhone hardware engineer Changlu. Apple says Tan solicited details about unreleased products in job interviews. Lou downloaded dozens of confidential hardware files and OpenAI was actually actively coaching departing employees on basically avoiding the kind of quote dreaded walk out when you leave that ends your access to confidential inter information. So uh at every level from members of its technical staff to its chief hardware officer this suit is saying and in coordination with business partners OpenAI has been stealing Apple's trade secrets and confidential information. Apple said, "So, Apple is seeking a jury trial. They want OpenAI to stop, destroy any proprietary materials, and redesign upcoming products to exclude Apple's technology.
1:00:48 · Open AAI's denied these allegations." And Paul, I mean, this is a pretty big bombshell. I mean, we're going to be watching this one closely. I did not see this one coming.
1:00:58 · Yeah. And just important context. So, you know, OpenAI acquired Johnny Ives company, right? Like $6 billion or something. Yeah.
1:01:07 · Um, so Johnny IV of of Apple fame the we've discussed like what are the products they're building. We've tried to kind of guess but we know it's hardware related and so we've known for a while that they were working on devices whether it's trying to compete with or replace the iPhone or if it's an ambient listening device that sits on your desktop or if it's a pen or a pendant like we don't we don't know yet.
1:01:29 · Um, but they are very aggressively moving into this and I would assume hardware is going to be a key part of OpenAI's IPO like the potential value and market opportunity behind their hardware.
1:01:39 · The um I'm not a lawyer but holy like the [laughter] the the lawsuit like what they're claiming it is it does not sound good for OpenAI. Um, so I I'll just uh a couple excerpts from Alex Heath who was uh I follow on on X and he was sort of reviewing the complaint. Um so he said yeah recent rec recently significant evidence has emerged suggesting this is quoting from the the filing individuals employed by open ad wrongfully took Apple secret and confidential information regarding our unreleased technologies processes and products. We always defend our team's hard work and innovations and we are taking all appropriate steps to do so.
1:02:18 · Um regarding an exapp employee named a lawsuit over several weeks while developing hardware for OpenAI Mr. L serendipitous or a ser serip wait how is that word surreptitiously I think there yes accessed and downloaded dozens of Apple's confidential hardware related files including uh voluminous detailed information about unreleased products engineering presentations technical specifications and proprietary project data other former Apple employees who had gone to work for open eye emailed themselves Apple's confidential information to personal accounts on their way out the door um regarding Tan a veteran Apple product leader who is Now, OpenAI's head of hardware. Apple's investigation has revealed Mr. Tan has been methodically using Apple's confidential information to benefit OpenAI. He has used an Apple internal project code name to ask quote, "What's the plan for X for an announced Apple product?" He has directed job candidates still working for Apple to bring actual parts from Apple to their interviews for showand tell sessions in which he and his team at OpenAI can elicit still more Apple confidential information. OpenAI has been instructing Apple employees to bring CAD design artifacts and prototypes to the interviews. Um, in February, an investigation was in the early stages. Apple wrote OpenAI to raise its concerns that Apple's confidential information could be making its way into OpenAI's business improperly. Apple asked OpenAI to discuss what precautions they were taking to avoid this problem, to investigate, and to uh remediate any issues. OpenAI did not respond. Um, OpenAI has been stealing Apple's trade tickets and confident information. As a result, OpenAI's naent hardware business now rests on the shakiest of foundations, rotten to its core by its illegal reliance on misappropriated trade secrets. [laughter] Um, February 9th, weeks after he had left Apple, knowing he had no right to do so, Mr. Lou tried to access Apple's network storage. He discovered that surprisingly he could still access the Apple network repository after leaving Apple, the result of then unknown authentication vulnerability. Rather than bringing this Apple's attention, Mr. Lou celebrated his fine with Miss Pang and said about exploiting it, quote, "lol, I found I can access the network storage so funny that's not going to play well in court."
1:04:29 · So, it goes on like it is it's insane.
1:04:32 · And like again, not an attorney, but like I just this is a bad bad look for OpenAI. Um, now OpenAI's statement, which I thought was actually a joke when I saw this. I was like, but this is from their director of strategic communications. He tweeted, "Our statement in response to this suit, quote, we have no interest in other companies trade tickets. We remain focused on building innovative technology that empowers people everywhere." I didn't realize that was the formal response from OpenAI. I was like, "Oh my god." Like their lawyers are going to tell them to get this down in like seconds, but apparently that was their actual tweet. And then Sam said, "I'm not afraid of Apple, but I have tremendous respect for them." Uh, okay Sam I again not sure that's going to go so well. Um, so my big questions are Apple and Open are in a partnership like they they're sharing chatt you can still use it in Siri. So that's probably not going so well. um what this does to Apple's or or OpenAI's hardware plans like I um and then the impact on the IPO. I I don't think when you're going for a multi-t trillion dollar IPO you want a massive lawsuit from Apple of all companies hanging over you. So I I don't know like this is really really interesting and a whole new thing to watch. To our point before about worrying that the model companies are taking your alpha, I would imagine at Apple you are not allowed to use chat GPT today. And if you still are, you shouldn't be. [laughter] It's kind of crazy.
1:06:01 · Yeah.
1:06:01 · Oh boy. Well, and that was what led to the Sam Elon tweet that I started with about like him stealing stuff. So, [laughter] well, I'm sure we'll have some updates here soon enough on that.
1:06:13 · Crazy.
Illinois Signs Nation-Leading AI Safety Law
1:06:14 · Next up, this past week, Illinois Governor JB Pritsker signed the Artificial Intelligence Safety Measures Act. This is a bipartisan law state leaders are calling the strongest AI safety and accountability framework in the country. This law targets the most capable models built by the largest companies using basically two thresholds. Um, if they have $500 million in annual revenue, they also kind of measure if they have massive computing that they're using for the models. And anyone who's covered by this must publish a transparency framework explaining how they apply industry standards, how they measure model capabilities and catastrophic risks, and identify and respond to safety incidents. This law also creates confidential reporting channels and whistleblower protections for any employees who raise safety concerns. So, Illinois actually then becomes the first state here to require regular independent third-party safety audits of covered AI systems. that actually goes a step beyond the 2025 New York and California laws that Illinois legislators used as a model. Illinois Attorney General Kwaame Raul will have authority to find companies up to $1 million for a first violation, up to 3 million for each additional violation.
1:07:27 · This law passed the Illinois House 110 to zero and it actually did draw support from OpenAI and Anthropic. Uh, Anthropic actually said it was proud to be the first AI lab to support the bill. So, Paul, I'm curious what you think of this law. Like, we're starting to see, it seems, states fill the void, which is kind of what the federal government, at least the administration was worried about. If there's no federal overarching law, states are going to pass their own regulatory frameworks. It sounds like, yeah, I didn't I didn't see any response from the government on this one. uh the federal government. I I would imagine they're not like huge fans of this and they probably aren't fans that open anthropic are publicly supporting it.
1:08:05 · Um I I don't know if this is like a skeptic in me, but I just find this hard to believe this is going to have any real impact. You know, if open anthropic are like, "Yeah, this is great." Then it's like, "Okay, yeah, it's probably [laughter] doesn't have much teeth and I don't know if it's like that it's only like a $3 million violation for each additional violation." They're like, "Yeah, okay. That's like a rounding error." um or that it really just isn't that aggressive. I I I don't know. It just doesn't seem like it got a ton of push back from anybody, which tells me it probably isn't like a huge deal, but it is a sign of states making progress and filling that void as you said. But I I don't know that it's like, you know, this is a transformational thing when it comes to, you know, how the models behave and what they're going to do and um the power that they're going to have. So, I don't know. I I might be wrong on that, but it just doesn't seem like it got a lot of run.
1:08:54 · And that tells me it probably isn't like a massive deal yet.
1:08:58 · And like we talked about with the California laws, they were considering like good luck keeping a close eye on like compute levels and where the thresholds are. Like I don't even know how you begin to do that at the state level.
1:09:10 · Yeah.
1:09:10 · And again, like one of the one of the loopholes that I just seem it seems like we've realized is going to be an ongoing issue is all of these regulations are related to publicly released models.
1:09:25 · So the the the loophole is they don't have to release the most powerful models that the the labs can have more powerful models. They can give you know exclusive access to select companies to government. And so this doesn't cover the fact that the models are just going to keep getting smarter and like maybe we just start reducing who has access to them because it's too hard to release them publicly. So they just do these limited releases and then the general populace doesn't ever get the most powerful models. I don't know. But like that seems like a logical path that all the regulation could lead to.
1:09:58 · Yeah.
AI Safety Index: Nobody Gets an A
1:09:59 · All right. Next up, something else related to AI safety. This past week, the future of life institute published their summer 2026 edition of their twice yearly AI safety index, which convenes an independent panel of seven AI experts to grade nine leading AI companies across 37 safety indicators. Unfortunately, the top grade was a C++.
1:10:22 · Anthropic again held the top spot with that C plus. Open AAI slipped from C++ to a C-grade. Google DeepMind ranked third after that. The panel noted all three have weakened or dropped earlier pledges to halt development on their own if certain red lines came into view.
1:10:40 · They have also soften their resistance to military uses of their technology. So Meta actually improved slightly, climbed from a D to a D+. Uh so great, good for you. Elon Musk's XAI uh obviously just rebranded to space XAI. It fell to an F.
1:10:58 · it joined China's Deep Seek and France's Mistl at the bottom. Um the institute's co-founder and president Max Tegmark who we've talked about before said the failing grades spanning three continents show that this is a global problem. Um he bas Stuart Russell who was also a panelist said companies have backed away from earlier commitments to release new systems only with safety measures appropriate for their capability levels.
1:11:21 · Now they're planning to release them even if it's demonstrably unsafe to do so. Tagged Mark told Time magazine that a real race to the top on AI safety will take regulation, saying he is cautiously optimistic and pointing to the new the EU's AI act, new Chinese rules that are taking effect this month and a more risk conscious US administration. So Paul like no question Tagmark comes at this from a very specific view on AI safety but him Russell the other panelists are pretty big people in AI saying that the companies are not doing a good job it sounds like yeah I don't think that's surprising to anybody I mean certainly you know some of the labs do a better job of releasing information being a little bit more transparent what they're doing being more vocal about the risks related to what you know everybody's building but some of these labs are you know very intentionally not sharing this kind of information. Yeah. like so that it's I do I do think it's funny that like the D to the D+ is is really good [laughter] and um but again if you know I know some of our listeners are really concerned about the AI safety side of this and where does this all go and so this is just to for you to know that there is a report out there that like anything in politics or in AI there's there is bias in the future of life institute like what do they believe and what are they pushing like you always have to put the context of okay, who is publishing this?
1:12:48 · What are they traditionally focused on?
1:12:50 · What is their mission? Um, all that being known though, like it's it's important to to know these things exist and be able to go down this path if you want to go read this stuff. And I there's nothing in it that surprises me. Like we know that they're not really doing a great job in this area, but it does sort of quantify it in some ways.
1:13:09 · Yeah, it really does. All right, so this past week we also saw a cheating scandal at Brown University that became a viral case study in what AI is doing to higher education. So this is about an economics professor, Roberto Serrano. He basically allowed his students to take take-home exams in his very difficult welfare economics course. This was the first time he had done it this spring.
The AI Cheating Scandal at Brown
1:13:34 · Unfortunately, it was due to because in December they had a campus shooting that left a bunch of students anxious about being in classrooms. So he said, "Okay, we'll do the exam as a take-home this year." Interestingly, when he announced that, enrollment jumped from a typical class of under 30 students to 86 students. The take-home midterm came back with an average score of 96 out of 100. 40 students scored a perfect 100.
1:14:01 · The historical average in the course ranged between scores of 65 and 80 on this exam. And he basically noted this.
1:14:07 · He saw that many answers had this weird convoluted style. So he started running the exam questions through chat GPT and produced basically similar answers. So what he did is he said look we're moving our final exam to inperson and if the score distributions look the same, we'll hold the midterm uh grades where they are and you can take your your well-earned grades. If the score distributions look different, we're voiding the midterm. The moment he said this, 18 students dropped the course.
1:14:36 · Nine more skipped the final. 22 of those 27 had scored a perfect 100 on the midterm. Here's the thing. The students then took the final in person and the average score dropped from 96 to 48. And so basically, while he says he's not using AI detectors or anything, he's like, "Look, this is a pretty solid indicator. Everyone's using chat GPT or other AI tools to basically totally cheat on all this stuff." And so it's kind of set off this scandal where the teacher's like, "Look, you know, he's softened the blow a little bit by changing some of the worst scores it seemed or given a little more credit, but he's like we've got a serious problem here." And so Paul, I don't know if this is surprising per se, but seeing those exact numbers was crazy to me.
1:15:22 · Yeah, that's the it's jarring just to to see like it quantified like this. I don't think anybody would be surprised that kids are using these tools to um so like this is a rapid fire I won't go deep on this but like the thing that jumps out to me is you can't outsource thinking and we have to as employers as parents um for ourselves for co-workers we have to be constantly aware that we have access to intelligence on demand and if we use it as a crutch for everything we will forget how to think And so this becomes critically important obviously for educators that they're they're not allowing they want to teach them how to use AI responsibly but you cannot do it where they just outsource the important part of what we do. Um and you have to think about this when hiring and when evaluating your existing employees. So in hiring you almost have to assume that the people you're interviewing out of college got through college using AI. Now, that can be really good, but you have to be able in your interview process to test for critical thinking capability to to make sure that they didn't um atrophy that ability in college that they've lost the ability to think for themselves and to assess things. So it's really really important like this is the whole next generation of workers will have never known education at the higher levels where they didn't have the ability to do it and they probably didn't have someone over their shoulder making sure they weren't using it um to to replace critical thinking. So that could be a problem that could compound within a company.
1:17:06 · Yeah, this feels like the real danger at the heart of all the positives about AI usage for students especially. It's like we're seeing it in the corporate world with what we've called in the past like work slop, right? Where people are just submitting AI responses to their co-workers that are just nonsense or like don't have any thought put into them and no makes more work for everybody. And it's like I just if you're a student and you think that you can come into a job and just submit whatever chat GBT gives you, I don't know if you people think that, but I would just say you have another thing coming if you think that is the standard moving forward. That's for sure.
1:17:41 · [clears throat] All right. So, another big story this past week. The New York Times reported that state actors in China, Russia, and to a lesser extent, Iran are working to inflame American debate over AI data centers. According to some analysis they're reporting on from the threat intelligence firm Althia. Between January and June, state media and the three countries mentioned data centers roughly 700 times, an average of nearly four times a day. basically in an effort to turn data centers and the controversy around them into what Althia calls quote a domestic fracture point heading into midterm elections where AI is seen as a top issue. So there were some examples that included a Chinese state-owned newspaper publishing a satellite image of a data center in Gainesville, Virginia warning that AI threatens Americans well-being. There is a chat GPT generated comic strip that was misinformation disguised as a Maryland news outlet blaming data centers for soaring electricity bills and a known covert Russian operation circulating a video attacking an American company's data center project in Armenia. So Paul, we have known for over a decade at least that foreign adversaries of the US have routinely pushed misinformation about tons of different subjects over the years to basically stoke domestic turmoil. Um, it seems to be happening with data centers. I believe this is yet another podcast prediction that came true or was correct. Um, is that this was happening. It seems like we have evidence of at least a little of it.
China and Russia Stoke US Public Opinion on Data Centers
1:19:10 · It's not something where it's like every single thing you see. There's a real and valid debate around the issue, but it sounds like some people are trying to take advantage of that.
1:19:19 · Totally. Yeah. I I was like I think when I first mentioned it like a month or so ago, I was trying not to like go down this rabbit hole too much. I mentioned like Cambridge Analytica and things like that. Um yeah, this is absolutely a strategy that foreign governments use to influence um citizens in other countries. The US does it to other countries and people do it to us. And so if you find something that causes friction um within the citizenship, then you you push on that and you use social media and now you can use AI tools to personalize the stuff that you know pushes these ideas. So, it's not to say, as you alluded to, Mike, that this isn't an issue, that there aren't real issues, but foreign governments are absolutely going to push on these things to try and create um a more uh I guess hate and fear and distrust among American citizens. So, it it's just a like people need to be aware and maybe you need to like educate your kids like if your kids aren't aware how social media works and things that they're seeing on TikTok and YouTube and wherever else they get their information that you know sometimes it's not real and it's actually meant to piss them off and get them all worked up about a topic. Uh AI is just one example of this, but like your kids need to know how the world works um for better or for worse. I've had this conversation with my 14 and 13year-olds. So, I don't know what when's too young to start them. And my kids aren't even on social media. Um they don't have accounts like they they use YouTube, but um they're not on Tik Tok and Instagram and things like that.
1:20:53 · And we've already had deep conversations about how this stuff works.
AI Use Case Spotlight
1:20:57 · Yeah.
1:20:57 · Well, okay. Next up, we have our AI use case spotlight where every week we give you a quick look under the hood at some real AI use cases we're exploring. So, I was just going to share one quick thing, Paul, that came out of my prep for our AI productivity workshop at our boot camp this week. Um, so as part of that, I was building out categories of AI capabilities basically to help participants better understand what can AI tools do today that you might not be aware of. Um, so I actually used codeex to build a huge sourcebacked like capability spreadsheet across chat GPT, Gemini, Claude and Microsoft Copilot. So basically the way I did this is like described what I wanted to do ideally and codeex split this research across four separate agents one per platform.
1:21:43 · So one agent research chat GPT using only official open AI sources. That was a specification. Same thing for claude Gemini and co-pilot. Each agent was asked to inventory the current enduser and business capabilities of each platform as of July. At this point it was July 8th or 9th 2026. So things like the models, the reasoning, chat, search, web grounding, deep research, etc. Like what are all the things that are today available in these tools? And it created it took like over an hour. I think this is like probably one of the longer use cases I think I've had for in a while.
1:22:20 · The final CSV has 877 rows. Do each one documenting a feature, showing what link and the documentation it came from. So there's like 200 roughly for each of these. We then did another pass to audit it. Basically looking for blank fields, looking for missing source links, stress testing all the answers. Super interesting. Um it's deeply overwhelming. It's not like useful as a public facing asset, I would say, but it was extremely critical in distilling this topic down into something others could understand and and making sure that was all backed by real data. So it was awesome. That was really cool. Uh I did a I shared this with you, Mike, so you've seen this, but I actually uh was trying to think of something to test Fable 5 with since they extended access to the 12th.
1:23:10 · And so I think it was like Friday morning or something or I guess no, Saturday. I'm looking at it now. I ran this on Saturday morning. So, uh I'm not a huge one on like tracking competitors.
1:23:19 · I don't I don't I generally just like focus on listening to our audience, looking at the trend data, and like building what needs to be built, what we think, you know, is best. But every once in a while it's good to just like see what's out there and see what's going on. So I actually ran a competitive analysis and I gave this to Chat GBT 5.6 soul high work edition it looks like and then I also ran it on Fable 5. So I'll give you the the exact prompt. It said run a competitive analysis on and then the name of the competitor. Consider strengths, weaknesses, threats and opportunities in comparison to our business. propose business strategies that we can use to exploit their weaknesses and our strengths to differentiate in the market and be uh the clear choice for enterprises. So that that was it. That's the entire prompt and then I gave it to both of them and it it crushed it like they were you saw Mike like like I and I what I did in this case again like the whole idea of AI slop I shared them internally with a couple people and I said listen this is I don't have time to edit this.
1:24:14 · I just wanted to run this as a a test in these two different models. I've taken both the outputs, I've put them into this single Google doc. Um, this is the raw unedited version of this. I'm not going to get to it until later next week, but just wanted you guys to have access to this as well. And so, again, like I think one, just the simplicity of the prompt. Two, using it for highle strategy. Three, if you're going to present something that you haven't verified yourself and given the time to think about, um, don't present it to your co-workers as though here's this genius thing I did. It's like, no, here's something I took 35 seconds to do across two platforms. Here is how I'm going to move forward verifying this and using this information, but here's the raw files in case you guys have a chance to take a look at it or it's relevant to anything else you're working on. So, um, that was a cool thing. I'm really anxious to actually dive into that later this week.
1:25:07 · Yeah, the outputs were so cool. And I would just mention as you're kind of talking about that, it really reminds me of what we talked about at the top of the episode that Ethan Mllik was saying that this kind of strategic knowledge work is not de is not coding. Um, so like you could spin up a bunch of agents to like stress test these ideas. It might result in something really useful, but you would need to then go audit like what the logic was there, right? So it' be much more you working in tandem with this raw output versus hey let me turn an agent loose and have it debug this thing. Right. So it's kind of very interesting to see how this knowledge work differs from coding.
1:25:44 · Yeah.
1:25:44 · And the other thing is you can then take this and say okay you know what I'm really intrigued by what they're doing. Uh every Friday morning run an update report. Tell me anything new they've done. And now you can get at the agentic side and start again changes work like it you reimagine what it's like to do these things like Mike when we were in my like the agency days you and I I mean we would charge like I don't know 10 to$20,000 to do these like deep competitive analyses and then provide strategy on top of it and it did it in 28 seconds. [laughter] It's wild.
1:26:17 · It's incredible.
1:26:18 · Yeah.
AI Product and Funding Updates
1:26:19 · All right. We're going to wrap up here with some AI product and funding updates. I'm going to run through it real quick because we have a ton of them. So, first up, OpenAI published an approach to government and national security partnerships, laying out principles for its growing defense and public sector work, including commitments not to allow its technology to be used for mass domestic surveillance, direct autonomous systems, or make high stakes automated decisions.
1:26:43 · OpenAI CEO of applications Fiji Simo announced she is actually stepping out of her role. She's stepping back to a part-time advisor role. She's been on medical leave for a chronic illness for 3 months, so she will no longer be in that role. Anthropic has, as of right now, once again, extended access to Fable 5 on all paid plans through Sunday, July 19th. It was originally supposed to end July 12th. Um, but now anyone who has a paid plan can use Fable 5 within certain limits rather than only paying us for usage as you go. So, when you're listening to this, you have a few more days to try it out.
1:27:20 · Anthropic appointed former Federal Reserve chair Ben Bernaki to its long-term benefit trust, the independent governance body that oversees the company's public benefit mission and can appoint board members. They cited how his ex his expertise on how AI will affect workforces and economies.
1:27:38 · Anthropic also brought Claude Co-work, its agent platform for delegating everyday knowledge work to web and mobile, letting users hand off a task at their desk, monitor progress from their phone, and let work keep running in the background. Microsoft has started replacing OpenAI and anthropic models with its own internally built MAI models for some AI features in Excel and Outlook. Meta released Muse Spark 1.1 an A Gentic encoding model and began charging for access to its models for the first time through the new Meta model API. Um pricing that CEO Mark Zuckerberg says the pricing is actually roughly a quarter of what anthropic and open AI charge. Funnily enough, he announced this in his first post on X since 2023.
1:28:26 · Meta also launched Muse image and Muse video. These are new image and video generation models. They immediately drew privacy backlash after the New York Times reported that public Instagram accounts are enrolled by default in a feature that lets other users generate AI images from their photos. The Financial Times also reported that Meta is testing prototype AI glasses designed for continuous ambient recording that feeds an onboard AI memory users can query. Uh Meta has reportedly considered whether the indicator light would actually stay on during passive data collection. SpaceXI, the new merge company formerly known as XAI and SpaceX is rolling out Grock 4.5. This is the new their new frontier model built for work beyond software engineering.
1:29:14 · They're rolling this out to all customers after a beta test program and cursor is making the model available across its coding platform. Amazon is working on a secret project cenamed Moonrakaker to turn Alexa into an AI agent that can chain together multi-step tasks from a single command with internal documents projecting more than hund00 million in GPU costs in 2026 alone.
1:29:40 · A startup called Prime Intellect raised a $130 million series A led by Radical Ventures with backers including Nvidia and Intel to build what it calls the open super intelligence stack. And artificial analysis introduced six new capability indices for comparing AI model capabilities across industry domains. And I think Paul you had one or two more things to highlight as well.
1:30:06 · Yeah, just this one just dropped this morning. So, I'll just throw this in here. So, um this is from the Stamford Digital Economy Lab. There's a new website. We must act now.I is the URL. 16 noble laureates join leading economists and AI researchers and call to prepare for AI's economic transformation.
1:30:30 · um calling for urgent preparation for the economic impacts of radically more powerful AI. This is led by Eric Brinolen, AJ Agarall, Anton Cornick, and Tom Cunningham. Um but it is also signed by Noam Brown who we talked about, Jeff Dean, Jack Clark, co-founder of Anthropic, Shoto Douglas, we've mentioned, Yasha Benjo, Eric Schmidt, Dean Ball, um Ben Bernaki, who you just mentioned. [clears throat] So the statement is three, it's three simple points. Um, one, AI may become radically more powerful uh over the next 10 years. Two, this could drive an unprecedented transformation of our economy, larger than the industrial revolution, but unfolding over a vastly shorter time frame. It could bring risks, including large-scale job displacement, as well as opportunities such as major gains in living standards.
1:31:20 · and three economists, policy makers and technology leaders must act now to understand the economics of transformative AI and to build the incentives, guardrails and institutions needed to steer AI in a direction that complements humans and benefits society.
1:31:36 · Seems like at least a few economists are waking up to what we have been talking.
1:31:41 · Yeah.
1:31:41 · Yeah. And it's interesting because it's like there's we mentioned recently how the labs sort of apparently coordinated effort all of a sudden started backing off of their belief like Sam Alman in particular even just tweeted again on Sunday like oh yeah I might have been wrong about jobs. It's like no you weren't like it's coming and like they're all you know starting to accept it. And I I think I don't know what the other goals behind this are. We'll put a link to the announcement from Stanford and a link to the site, but uh it's it has to happen.
1:32:12 · Like I've always said, at minimum, we need to be prepared. If it doesn't happen, great. But we need to be prepared in the event that it does cause massive displacement and uh underemployment.
1:32:25 · All right, so that's a wrap for this week, Paul. Just one more quick reminder to visit our AI poll survey for this week, smartrx.ai/pulse. This week we are asking about your feelings around chatbt work and also your feelings around using voice AI. So if you go to that link smarterx.ai/pulse, we'd love if you shared your thoughts.
1:32:47 · Paul, thanks again.
1:32:48 · Yeah.
1:32:48 · Programming note, no episode uh July 21st. So, right. So, I will be on vacation and we were going to try and squeeze in a recording early, but Mike and I are in Columbus uh for the AI for business boot camp this week. So, it just it wasn't going to work despite our best effort. So, we're going to one week break for the podcast. We'll call it our summer break. Mike, give you an day off.
1:33:13 · There you go.
1:33:14 · And uh and then we will be back on I guess that would be what? July 28th.
1:33:19 · Yep.
1:33:19 · Yep.
1:33:19 · All right. So, thanks everyone.
1:33:21 · Have a great two weeks, I guess. And uh we will be back with the latest on episode 226.
1:33:27 · Thanks for listening to the Artificial Intelligence Show. Visit smarterx.ai to continue on your AI learning journey and join more than [music] 100,000 professionals and business leaders who have subscribed to our weekly newsletters, downloaded AI blueprints, attended virtual and [music] in-person events, taken online AI courses, and earned professional certificates from our AI Academy, and engaged in the Smarter X Slack community. [music] Until next time, stay curious and explore AI.