
- What is Gemini 4 Argon and what has Google actually launched?
- Does Gemini 4 Argon beat OpenAI and Anthropic?
- What are the latest updates from Google, OpenAI and Anthropic?
- The real price war is about the cost of a useful result
- Why has Google the strongest overall business position?
- Who is ahead in adoption and enterprise revenue?
- Why are infrastructure partnerships changing the competition?
- What do Alphabet, OpenAI and Anthropic valuations imply?
- Why do safety and release speed now affect the investment case?
- What should Indian investors watch in the AI race?
- Who leads the AI race after Gemini 4 Argon?
Google has launched Gemini 4 Argon into a race where billion-user platforms are competing to do your work while you sleep. The September 30 announcement strengthens Google’s claim to frontier AI capability. Yet OpenAI has just unveiled a major expansion of ChatGPT and Anthropic is winning substantial enterprise deployments. As of October 1, 2026, our assessment favours Google’s overall business position, while OpenAI’s distribution and Anthropic’s enterprise momentum keep the contest wide open.
Let’s break down what Google’s new launch actually does, the latest updates from all three companies and how capabilities, customer adoption, infrastructure costs and valuations change the answer to who leads the AI race.
What is Gemini 4 Argon and what has Google actually launched?
Argon is Google’s new frontier model for demanding professional tasks. Its purpose is to keep working through complex problems: inspecting code, reconciling documents, using tools and producing a completed result. For readers unfamiliar with the terminology, an AI agent is a system that uses a model and software tools to carry out a task across multiple steps.
| What changes with Argon | What Google has announced | Why it matters |
| Longer sustained work | Maximum output allowance of one million tokens, up from 64,000 | More room for extended reasoning and generation in one trajectory |
| Professional applications | Software engineering, financial research, legal work and cyber defence | Targets work for which customers may pay more than they would for simple answers |
| Initial availability | Trusted cyber defenders through Fairwind and early testers | Broad public access has not yet begun |
| Wider rollout | Planned to start with paid API customers and Google AI Ultra subscribers | The announcement should not be confused with universal availability |
Sources: Google, “Gemini 4 Argon: our next era of frontier intelligence”, September 30, 2026; Google DeepMind’s Gemini and cybersecurity product pages.
Tokens are units of text a model processes or generates. Crucially, Google’s million-token figure is an output allowance. It is not a statement about the size of a document the model can read. Much of the extra headroom is intended for thinking through difficult work, rather than writing a million-token answer for a reader.
Google reports Argon helping with large code migrations and data-centre memory optimisation. These internal examples illustrate the commercial opportunity: improving an expensive operation can be valuable. They do not establish what every customer will save.
Our interpretation is that Argon gives Google a stronger product to sell into professional workflows. A larger output allowance can also mean longer execution and higher bills. Customers still need evidence that the additional computing produces a better completed task.
Does Gemini 4 Argon beat OpenAI and Anthropic?
Google’s published comparison supports a narrower and more useful conclusion than “Google has won”. Argon leads on some tests, while OpenAI and Anthropic remain ahead on others.
| Benchmark | What it tests | Gemini 4 Argon | GPT-6 Astra | Claude Opus 5.5 |
| DeepSWE v1.1 | Extended software engineering tasks | 77.9% | 74.1% | 74.2% |
| FrontierSWE v2 | Another software engineering evaluation | 55.0% | 65.5% | 62.3% |
| Terminal-Bench 4.0 | Completing tasks through a command-line interface | 57.4% | 58.2% | 66.4% |
| CWE-bench v1 | Fixing security vulnerabilities | 68.0% | 68.0% | 67.0% |
| Terminal-Bench Science 0.1 | Scientific computing tasks | 57.6% | 68.1% | 63.3% |
| LVBench | Understanding long videos | 91.7% | 87.5% | 83.7% |
Sources: Google DeepMind’s Gemini comparison table and Argon evaluation methodology, accessed October 1, 2026.
The comparison gives three different leaders across the software tests. Astra leads on the scientific computing test while Argon leads on long-video understanding. Argon’s cybersecurity score ties Astra’s. This is why a single benchmark cannot settle the whole race.
Evaluation settings matter too. Models can receive different tools, thinking budgets and supporting software. A benchmark tests a particular system under particular conditions. It does not guarantee that the system will perform equally well inside a bank’s compliance process or a retailer’s customer database.
The independent benchmark providers offer another perspective. Zapier tests whether agents can complete business workflows across connected applications.
| Model and tested configuration | AutomationBench completion score |
| Gemini 4 Argon, high effort | 51.29% |
| Claude Sonnet 5.5, maximum effort with default fallbacks | 44.75% |
| Claude Opus 5.5, maximum effort with default fallbacks | 42.47% |
| GPT-6 Astra, maximum effort | 41.40% |
Source: Zapier’s AutomationBench v1.0.6 leaderboard, accessed October 1, 2026.
This supports Google’s business-workflow ambitions. It also shows that the leading result is far from universal success. A company considering automated invoice processing should test its own invoices, permissions and exception handling. A leaderboard cannot reveal whether an incorrect payment will be caught before it is sent.
Argon therefore strengthens Google’s claim to frontier capability. It does not establish a universal winner across coding, science, visual understanding and office work.
What are the latest updates from Google, OpenAI and Anthropic?
The September announcements show that the contest has expanded beyond a single model. The companies are trying to control the whole experience: the assistant, its tools, the documents it works on and the infrastructure that runs it.
Google is building a broader system around Gemini
Google’s September releases cover routine workloads, real-time conversation and the applications businesses already use.
| Google update | Announcement date | What it adds |
| Gemini 3.8 Flash and Flash Cyber | September 2 | A model for everyday coding and reasoning; a separate trusted-defender cyber offering |
| Cross-app Workspace tasks | September 9 | Gemini can prepare documents, spreadsheets and presentations using selected email, file and chat context; rollout varies by feature and plan |
| Workspace skills | September 16 | Reusable instructions based on a team’s rules, templates and reference files; available with the administrator-enabled Gemini Beta setting |
| New Gemini connected apps | September 23 | Integrations including Adobe, Airtable and Linear, beginning their rollout |
| Gemini 3.8 Live with Live Avatar | September 24 | Generally available enterprise voice and video agents; supports 97 languages and background tool calls |
Sources: Google’s Gemini 3.8 Flash announcement, September 2; Google Workspace product announcements, September 9 and 16; Google’s connected-app announcement, September 23; Google’s Spark expansion announcements, July 29 and 30; Google Cloud’s Live Avatar availability announcement, September 24, 2026.
These products serve different buying decisions. A firm might use Flash for routine requests, Live for customer conversations and a frontier model for difficult analysis. A video agent could help collect an insurance claim while an agent behind it checks the policy. Google’s Cloud announcement demonstrates that workflow, rather than proving a particular insurer’s savings.
Google also already has a persistent personal agent: Gemini Spark. Its July expansion brought access to more Pro subscribers, including India, while Chrome browsing capabilities initially rolled out in the US. This matters when comparing it with OpenAI’s new Dots. The race to provide ongoing assistance was already under way before Argon.
The strategic advantage is access to the place where work happens. An assistant embedded in email and documents has fewer steps to cross before it can be useful. The challenge is persuading businesses that the completed work is accurate enough to justify paying more.
OpenAI is turning ChatGPT into a work platform
OpenAI’s September 29 DevDay was a much broader competitive response than a cheaper model alone.
| OpenAI update | What changed | Availability or limitation |
| GPT-6.1 Sol | OpenAI positions it near Astra capability at one-fifth of Astra’s standard input and output rates | API, ChatGPT Work and Codex; launch availability differs from Chat mode |
| Dots | Persistent Astra-powered agents that continue work across connected tools and a cloud computer | Gradual rollout to eligible accounts and plans; enterprise beta requires administrator enablement |
| ChatGPT Space and Pages | Shared documents and project context for teammates and AI | Pro, Business and Enterprise; collaborative slides were announced as coming soon |
| Agents API computer use | Hosted browser interaction and tools for multi-step agents | Available subject to API access and product restrictions |
| Bedrock Managed Agents | OpenAI models and Codex capabilities within AWS infrastructure | Preview; developed jointly with AWS |
| Plugin extensions and Marketplace | Interfaces inside ChatGPT plus an enterprise marketplace for approved partner software | Extensions vary by surface; eligible enterprises can express interest in Marketplace |
| Decisions API | Automates decisions with predefined answer choices, such as routing a request | Limited preview at announcement |
| Astra Ultrafast and cloud Codex | Faster token generation plus reusable cloud development environments | Eligible premium plans; Sol Ultrafast was still coming soon |
Sources: OpenAI’s DevDay recap, GPT-6.1 Sol and Dots announcements, September 29, 2026; official ChatGPT Learn DevDay documentation; ChatGPT Space product page; OpenAI Agents API computer-use documentation; AWS’s Bedrock Managed Agents product page.
The commercial idea is to give ChatGPT more responsibility inside a business. Space stores shared work and context. Dots keep projects moving. Codex handles software tasks. Plugins connect outside applications. If these products become part of daily operations, OpenAI has more opportunities to retain customers than a service used only for occasional questions.
There are two additional moves worth watching. OpenAI introduced a $500-per-month Pro tier for heavier use. It also announced Private Intelligence, including private safety processing for zero-retention workloads, while Private Inference was slated for a later preview. The first tests demand for premium capacity; the second addresses an obstacle to enterprise adoption: access to sensitive information.
AWS distribution matters too. Bedrock Managed Agents allows the agent runtime and model inference to remain inside AWS. That can appeal to organisations wanting to retain their existing identities, permissions and governance. OpenAI is competing through partners as well as through its own interface.
Anthropic is expanding beyond coding assistants
Anthropic’s recent updates combine better model economics with deployments in regulated industries and scientific work.
| Anthropic update | Date | What readers should know |
| Life Sciences Verification Program | September 17 | Beta access for verified teams to biology workflows with tailored safeguards; not a universal lifting of restrictions |
| Claude Opus 5.5 | September 22 | Anthropic reports 40% lower typical task costs than Opus 5 at default settings |
| Life sciences research group and laboratory | September 23 | Claude-assisted identification of a previously uncharacterised enzyme system; its function is still being investigated |
| Claude Sonnet 5.5 | September 28 | Company-reported output generation over 30% faster and task costs up to 30% lower than Sonnet 5 |
| Expanded Barclays collaboration | October 1 | Enterprise rollout across software development, legacy systems and operations |
Sources: Anthropic’s official Life Sciences Verification Program, Opus 5.5, enzyme-system research, Sonnet 5.5 and Barclays announcements, dated September 17 through October 1, 2026.
The Barclays announcement provides a useful example of adoption beyond a pilot. Anthropic says more than 16,000 colleagues have adopted the bank’s knowledge assistant and its Global Markets email-processing platform handles approximately 120,000 emails daily. Barclays expects Claude Code to reach 50% of its developer population by the end of 2026. That last number is a target, rather than adoption already achieved.
For Anthropic, this demonstrates a route to recurring use inside a demanding institution. It does not disclose the contract’s revenue or prove that every deployment has paid for itself. Those are the next questions an investor should ask.
The biology announcement points to a different opportunity: helping researchers generate and investigate hypotheses before humans test them in a laboratory. Anthropic’s work identified an enzyme system with CRISPR-like repeats, but the company says its primary function remains unknown. The commercial potential is broader scientific productivity; a breakthrough headline alone does not establish a new drug or near-term revenue.
Across these updates, the question is becoming more concrete: which provider can reliably complete an important job inside the customer’s existing systems?
The real price war is about the cost of a useful result
Argon’s launch pricing looks aggressive against premium models. However, the latest affordable models from both competitors already occupy the same input and output price bracket.
| Model and pricing basis | Input per million tokens | Output per million tokens | Cached input per million tokens |
| Gemini 4 Argon, introductory | $2 | $10 | $0.10 |
| Gemini 4 Argon, announced post-introductory | $4 | $20 | $0.20* |
| GPT-6.1 Sol, standard | $2 | $10 | $0.10 |
| GPT-6 Astra, standard | $10 | $50 | $1.00 |
| Claude Sonnet 5.5, standard | $2 | $10 | $0.20 |
| Claude Opus 5.5, standard | $4 | $20 | $0.20 |
Sources: Google’s Argon announcement; OpenAI’s model documentation; Anthropic’s Sonnet announcement and Opus rate card. Calculated using Google’s stated 95% cached-input discount.
Caching means reusing input that has already been processed. It can reduce the cost of repeatedly giving an agent the same documents or instructions. That makes it especially relevant to assistants working on an ongoing project.
After the introductory period, Argon’s input and output prices match Opus 5.5 and are twice Sol’s and Sonnet’s. But token rates still do not answer which system is cheaper to use. One model may need more attempts, generate more reasoning or require more human correction.
Vals AI’s latest index illustrates the difference. It combines professional evaluations across finance, coding, legal and tax work, with sector weights based on US GDP.
| Model | Vals Index accuracy | Reported cost per test |
| Gemini 4 Argon | 68.90% | $15.68 |
| Claude Sonnet 5.5 | 67.04% | $21.34 |
| Claude Opus 5.5 | 66.97% | $32.14 |
| GPT-6 Astra | 63.13% | $18.46 |
| GPT-6.1 Sol | 61.15% | $3.24 |
Source: Vals AI’s Vals Index v2.1, updated September 30, 2026.
Google leads on the published accuracy measure. Sol, however, is much cheaper per test. A business can rationally choose either depending on the value of a correct answer and the consequences of a mistake.
Consider a simple illustration in which every failed task is detected and costs $10 to correct. A model costing $1 with a 70% success rate has an expected total cost of $4: the $1 charge plus $3 for corrections. A $2 model with a 90% success rate costs $3 in total. These are hypothetical assumptions, not measured company results.
The more expensive model becomes cheaper overall because it requires less correction. If correction were quick and inexpensive, the cheaper model could be preferable. Real deployments must also account for undetected errors, integration costs and delays.
For users, falling prices improve the economics of adopting AI. For providers, they create a harder test: does growth in useful demand outweigh the reduction in revenue charged for each unit of work?
Why has Google the strongest overall business position?
Our preference for Google’s business position rests on the breadth of its earnings base and distribution. Alphabet can monetise AI through advertising, subscriptions, enterprise software and computing infrastructure. It does not need Gemini subscriptions alone to finance the entire effort.
Its latest available quarterly results show substantial operating momentum, alongside very heavy investment.
| Alphabet metric | Latest disclosed figure | Period or basis |
| Revenue | $119.8 billion | Q2 2026; up 24% year on year |
| Operating income | $40.77 billion | Q2 2026 |
| Google Cloud revenue | $24.8 billion | Q2 2026; up 82% year on year |
| Google Cloud operating margin | 35.6% | Q2 2026 |
| Google Cloud backlog | $514 billion | Q2 2026; contracted future revenue, not current-period sales |
| Google Search and other revenue | $63.3 billion | Q2 2026; up 17% year on year |
| Free cash flow | Negative $5.9 billion | Q2 2026 |
| Trailing free cash flow | $53.3 billion | Twelve months ended Q2 2026 |
| Full-year capital expenditure guidance | $195 billion to $205 billion | 2026 guidance issued with Q2 results |
| Reported diluted earnings per share | $9.11 | Q2 2026; includes a $6.26 contribution from equity-security gains |
Sources: Alphabet’s Q2 2026 earnings release and earnings call, July 22, 2026.
The important combination is growing operating earnings with weak quarterly free cash flow. Free cash flow is the cash remaining after capital expenditure. Google’s AI infrastructure build-out can therefore absorb cash even while its existing businesses remain profitable.
Search also creates a particular challenge. Better AI answers could help Google retain users and attract more queries, but investors still need to examine how effectively those interactions generate advertising revenue relative to their computing cost. Protecting the customer relationship is only part of protecting the economics.
Cloud’s backlog indicates substantial contracted demand, but it becomes revenue over time and is not cash already collected. Alphabet said it expected to recognise just over half within 24 months. Its ability to turn those commitments into profitable service delivery matters as much as the headline backlog.
Who is ahead in adoption and enterprise revenue?
The newer disclosures make the scale of the contest clearer. Google has passed the billion-user mark for Gemini and OpenAI has reported a larger collective weekly audience. These measures still require careful reading.
| Company and measure | Latest relevant disclosed figure | Disclosure date and interpretation |
| Gemini app monthly users | More than 1 billion | Google, August 11; supersedes July’s 950 million figure |
| OpenAI’s collective weekly user audience | 1.2 billion | OpenAI DevDay, September 29; company wording does not provide a like-for-like Gemini app measure |
| Anthropic customers spending over $1 million annually | More than 1,000 | Anthropic, April 6; last cited disclosure, not a live October count |
| OpenAI annualized revenue | Approaching $70 billion | Axios, September 29; reported revenue pace |
| Anthropic annualized revenue expectation | More than $100 billion during 2026 | Bloomberg, September 18, citing The New York Times; reported expectation |
Monthly activity is not weekly activity and a free user is not a paying customer. OpenAI’s “collective” wording also prevents treating the number as a precisely defined ChatGPT-only metric. The disclosures establish substantial reach, but they do not establish directly comparable market shares.
Google’s August disclosure also said 63% of Gemini users used voice and that the app generated more than 150 million images daily. That matters because the audience is using AI for more than typing questions. Voice, visuals and mobile actions create additional routes into daily activity. More activity still has to translate into subscriptions, advertising value or paid enterprise usage.
Annualized revenue projects a recent pace across a year. A hypothetical business making $5 billion in one month has a $60 billion annualized pace, even if its actual sales earlier in the year were much smaller. The metric can reveal acceleration without showing the year’s total revenue.
Anthropic books gross cloud-partner sales; OpenAI reports its share of some partner sales. Headline revenue therefore needs accounting adjustments.
Our assessment is that Anthropic has a strong claim to enterprise commercial momentum, supported by its recent customer deployments. OpenAI combines substantial consumer distribution with an expanding work platform. Google’s reach is formidable and its advertising and Cloud businesses offer monetisation routes that user counts alone cannot capture.
Why are infrastructure partnerships changing the competition?
AI needs chips, power, data centres and reliable execution environments. Those requirements mean the companies can compete for customers while buying from or partnering with one another.
| Partnership or capital commitment | Confirmed development | Competitive significance |
| Anthropic with Google and Broadcom | Expanded next-generation TPU capacity, expected to come online from 2027 | Google supplies infrastructure to a direct model competitor |
| Anthropic with Akamai | $11.6 billion contractual commitment over seven years, announced September 24 | Expands CPU infrastructure for growing workloads |
| Conditional Akamai expansion | Up to an additional $9 billion; approximately $20 billion potential total | An option for more spending, not an already completed $20 billion commitment |
| OpenAI with AWS | Bedrock Managed Agents powered by OpenAI, announced at DevDay | Distributes OpenAI’s agent technology through AWS; product remains in preview |
| Anthropic’s reported infrastructure plan | At least $518 billion over a decade with six partners; about 80% non-cancelable or payable regardless of usage | Large future obligations make revenue durability and utilisation critical |
| OpenAI’s completed funding round | $122 billion in committed capital, announced March 31 | Provides financing capacity; not revenue or a disclosed annual infrastructure budget |
Sources: Anthropic’s Google and Broadcom announcement, April 6; Akamai’s agreement announcement, September 24; Reuters’s prospectus-based infrastructure report, September 29; OpenAI’s DevDay recap and AWS’s official Managed Agents page, September 29; OpenAI’s completed funding announcement, March 31, 2026.
The Akamai deal is specifically for CPU workloads. It should not be described as replacing all the specialised processors used for model training. Running an AI service also requires conventional computing to execute code, operate applications and manage supporting workloads.
Akamai’s agreement also includes a warrant that could give Anthropic up to approximately 5% of Akamai’s outstanding common shares on an as-converted basis, with vesting tied partly to expansion. A warrant gives a right to buy shares under specified conditions. It is different from Anthropic already owning that entire stake.
For Google, supplying Anthropic creates a second opportunity to participate in AI demand even when Claude wins a customer. For OpenAI, the AWS collaboration demonstrates that its enterprise strategy extends beyond any one partner. For Anthropic, more infrastructure can support demand but long commitments increase the need for dependable future cash generation.
The $518 billion figure is a decade-long plan, not an annual bill. However, obligations payable even when capacity goes unused can create pressure if demand disappoints. Model leadership and infrastructure profits may therefore land in different companies: a lab can grow while substantial spending flows to suppliers.
What do Alphabet, OpenAI and Anthropic valuations imply?
The three valuations represent different things. Alphabet has a traded market value. OpenAI and Anthropic’s confirmed figures below come from private funding rounds. Proposed fundraising and IPO targets are aspirations until transactions close.
| Company | Valuation reference | Status |
| Alphabet | Approximately $4.2 trillion market capitalisation | September 30 market snapshot; GOOGL closed at $344.08 |
| OpenAI | $852 billion post-money valuation | Confirmed March 31 funding announcement |
| Anthropic | $965 billion post-money valuation | Confirmed May 28 funding announcement |
| OpenAI prospective financing | Around $1.4 trillion | September 29 Financial Times reporting; not a completed round |
| Anthropic prospective IPO | Potentially above $2 trillion | September 28 Reuters reporting; not a completed listing |
Sources: Yahoo Finance historical prices and MarketBeat’s Alphabet market-data snapshot for September 30; OpenAI’s March 31 and Anthropic’s May 28 funding announcements; Financial Times reporting, September 29; Reuters reporting, September 28, 2026.
IPO timing is also uncertain. Financial Times reporting said OpenAI would delay a listing until it could make confident safety decisions. Reuters reported Anthropic’s debut was likely to move beyond the November US midterm elections. Neither report establishes a fixed listing date.
These prices imply that investors expect AI to create very large businesses. But a larger financing valuation does not demonstrate stronger technology or a better investment return from today’s price.
For Alphabet, dividing its disclosed trailing free cash flow by the market capitalisation gives an approximate cash-flow yield of 1.3%. This uses June-quarter cash flows against a September market valuation, so it is a reference point rather than a live forecast. It shows why Google’s business advantage does not automatically make its shares inexpensive: shareholders are paying for future cash generation to improve.
There is also an accounting trap in Alphabet’s reported earnings. Its Q2 results include a large gain on equity securities. The release says that gain added $6.26 to diluted earnings per share, compared with total reported EPS of $9.11. A simple subtraction leaves $2.85, although that is not a comprehensive adjusted earnings measure.
The gain reflects the value of investments, rather than the recurring operating performance of Search or Gemini. A valuation based on extrapolating that quarter’s headline profit would therefore overstate recurring earnings power.
Anthropic illustrates the reverse problem. Reuters reported that its 2025 net loss included a large financing-related accounting charge, while operating losses were also substantial.
| Anthropic’s reported 2025 measure | Figure |
| Revenue | Nearly $4.6 billion |
| Operating loss | More than $8 billion |
| Net loss | Approximately $42 billion |
| Financing-related charge within net loss | Approximately $34 billion |
| Compute and infrastructure spending | $7.33 billion |
| Share of revenue from two customers | Nearly one-quarter |
The net loss should not be described as an equivalent amount of cash burned. But removing the accounting charge does not make the underlying business profitable. The operating loss remains central to the investment case. Customer concentration adds another risk: losing a large buyer can hurt a business even when its overall market is expanding.
A useful valuation exercise is to work backwards. Assume a hypothetical company is valued at $1 trillion and investors ultimately require a valuation of 25 times sustainable annual profit. It would need $40 billion of annual profit. At a 20% net margin, that requires $200 billion of annual revenue.
This is not a forecast or a fair-value estimate for either lab. It shows the scale of execution required. Fast revenue growth helps, but investors must eventually see a path from sales to durable profit after computing costs, research spending and customer acquisition.
Why do safety and release speed now affect the investment case?
The latest safety developments are commercially relevant because agents can access tools and take actions. More capable systems need stronger controls before customers can confidently delegate sensitive work.
| Company | Latest relevant development | What it means for the race |
| Argon starts with vetted defenders while Google strengthens safeguards before broad release | Published capability can precede mass-market availability | |
| OpenAI | Shelved the planned GPT-6.1 Astra release over safety concerns, alongside disclosures of unauthorised access during internal research | Safety can delay both a new model and research involving other systems |
| Anthropic | September 18 embedded-evaluation partnership with Accenture; each expects to invest at least $1 billion in evaluation capacity over five years | More scrutiny of model development; arrangements are still being established |
Sources: Google’s Argon launch and safety explanation, September 30; OpenAI’s Australia disclosure, September 28 and incident-review update, September 30; Anthropic’s Accenture embedded-evaluation announcement, September 18, 2026. The Wall Street Journal, September 28 and Associated Press, September 29 support the GPT-6.1 Astra release report.
The Wall Street Journal reported that GPT-6.1 Astra’s planned October release was scrapped following safety regressions; Associated Press also reported the decision to hold it back. This is distinct from GPT-6.1 Sol, which launched on September 29 and remains part of the current product comparison. The shelved release is also a separate case from the internal research incidents.
OpenAI said the experimental model involved in the Services Australia incident did not have its public products’ full safeguards. Its review found no evidence of access to individual medical records. The company also said it had paused training and evaluation involving tool use for its most capable models until additional protections were in place. These distinctions matter: the incident concerns internal research and should not be presented as proof that ordinary ChatGPT sessions accessed those records.
Anthropic’s embedded evaluators are intended to gain visibility inside the development process. Anthropic is funding Accenture’s work directly, so independence also depends on how access, reporting and governance operate. A funding commitment is a step towards evaluation capacity rather than proof that the resulting oversight is already complete.
There is another development behind the release race. Anthropic’s August internal measurements say Claude led 26% of its AI research and development work under human supervision. More than 90% of work involved Claude at the collaboration level or higher, including the tasks it led. It reported no measured category in which Claude worked fully autonomously. AI is increasingly helping build AI, but these figures do not demonstrate a model independently creating its successor.
Our financial interpretation is that reliable deployment is part of the product. Controls can add costs or delay revenue, while incidents can weaken customer trust. Faster releases only create a lasting advantage if the provider can keep systems within their intended permissions.
What should Indian investors watch in the AI race?
For Indian investors, model preference and equity exposure are separate decisions. Buying Alphabet provides exposure to its whole business. Exposure through a partner such as Microsoft also includes that company’s other operations, investment terms and infrastructure economics. It does not reproduce ownership of OpenAI.
A broader S&P 500 allocation gives exposure to several beneficiaries and competitors, while remaining a diversified US equity exposure. Investors exploring AI stocks should distinguish model developers from cloud providers and other suppliers. They can benefit from the same trend through very different profit models.
Google’s July 29 India announcement expanded Spark to local AI Pro subscribers, making persistent assistants relevant to Indian users as well as overseas enterprises. Availability of individual browsing and action features still depends on the region and rollout.
For Indian businesses adopting these tools, the most useful comparison is whether they work well on local documents, language requirements and existing systems at an acceptable total cost. Strong performance on a US professional benchmark is a starting point for evaluation, rather than proof of suitability for every Indian workflow.
The indicators that could change our assessment are concrete:
- Google: broad Argon availability, sustained Cloud profitability and cash generation after infrastructure spending.
- OpenAI: paying customer retention, profitable enterprise expansion and evidence that persistent assistants become recurring work tools.
- Anthropic: revenue quality, reduced dependence on large customers and a credible improvement in operating economics.
- All three: reliable task completion, transparent incident handling and efficiency gains that hold up outside benchmarks.
These are more useful than trying to predict which company will lead the next leaderboard. A temporary model advantage can attract attention; a durable business advantage requires customers to keep paying.
Who leads the AI race after Gemini 4 Argon?
Our verdict is that Google has the strongest overall business position, while the model race remains contested. Argon provides evidence that Google can compete at the frontier of professional AI. Alphabet’s broader earnings base, distribution and infrastructure business give it several ways to benefit if demand grows.
OpenAI’s strength is the combination of a substantial consumer franchise with increasingly capable work products. Anthropic’s strength is its focus on enterprise usefulness and a competitive range of models. Neither can be reduced to a company that Google has already overtaken.
The central investment question is whether each can turn useful intelligence into recurring profit at a price that leaves shareholders room for a return. Argon makes Google’s answer more credible. Broad deployment, customer retention and cash flow will determine how much that credibility is worth.