본문으로 바로가기메뉴 바로가기
Anthropic and OpenAI face an independence test as 100-plus evaluators speak out
공유

Anthropic and OpenAI face an independence test as 100-plus evaluators speak out

Investor Read-Through: Trust Becomes an AI-Control Variable

Anthropic and OpenAI now face a credibility test that could influence how investors assess the wider AI sector: whether outside evaluators can inspect frontier systems independently, transparently and without retaliation. On Sept. 18, the AI Evaluator Forum consortium organized a public letter signed by over 100 AI experts and evaluators calling for scientific objectivity, independence, transparency and robust protections for third-party safety work.

The immediate market implication is not a disclosed change in revenue, margins or product demand. It is a governance question that can shape access to capital, enterprise confidence and the policy environment around companies building foundation models. The letter's requirements are concrete, but implementation remains undefined: the exact number of signatories, which evaluators would be selected, how deeply they could inspect systems and whether companies will act are all unknown.

AD

What the AI Evaluator Forum Letter Actually Requests

Third-party evaluation means an organization outside the model developer examines both the systems and the development process. The coalition says credible embedded evaluations should cover model risks and significant real-world harm, as well as training, deployment, oversight, operational and safeguard practices. It also argues that evaluators need protection from interference by the companies being examined.

Conrad Stosz, chair of the AI Evaluator Forum, described the objective as establishing “a shared common ground on basic principles” so independent oversight can help manage AI risk. The coalition said it does not advocate for “one particular way” to make models safe; its request is for basic principles and greater standardization. One example cited in the letter is the A-EF-1 standard, although the group says further work is necessary before embedded evaluations are consistently effective and meaningful.

That distinction matters for investors. A standard can make reports more comparable, but it does not by itself determine what access a laboratory grants, what information becomes public or how a company responds to an adverse finding. The letter therefore creates a framework for accountability rather than a verified safety outcome.

Why Access Is the Core Dispute

Anthropic CEO Dario Amodei proposed giving some evaluators “employee-like access” to inspect and audit frontier models and development processes. In the scenario described by Stosz, evaluators could use company computers, speak candidly with employees, and see sensitive internal data and unreleased systems. That level of access could produce more information about systems being used internally rather than only models released to customers.

OpenAI CEO Sam Altman, SpaceX-associated Elon Musk and Microsoft CEO Satya Nadella have publicly supported Amodei's proposal. Support, however, has not resolved the operating questions: which evaluators qualify, what technologies they can inspect and how confidential information would be handled. The fact that these questions remain open is central to the investment read-through. A pledge without an access protocol is difficult to compare across companies or to price as a durable reduction in risk.

The source also identifies an unreleased OpenAI model in connection with an attack involving Hugging Face, but its specific identity is unknown. That limitation prevents investors from drawing conclusions about the model, the incident or any technical failure from the information available here.

Independence Versus Internal Control

The coalition's argument is that evaluators cannot be credible if the businesses they assess control their access or can retaliate against them. Vinh Nguyen, a Council on Foreign Relations senior fellow, former chief AI officer of the National Security Agency and a letter signatory, said the public and government should not depend solely on a few powerful labs' own account of what is secure when capabilities could affect cybersecurity, critical infrastructure, national security and the economy.

That position does not eliminate internal safety work. Stosz said third-party evaluators are not intended to “be a replacement for any internal efforts to evaluate.” The letter instead frames embedded reviews as a complement to broader external oversight, including public transparency and wider access for independent researchers.

For investors, the mechanism is straightforward but conditional. Stronger external scrutiny could improve confidence in risk disclosures and reduce the gap between a company's internal description of safety and what independent experts can verify. Conversely, intrusive access, unclear confidentiality rules or a limited pool of technically credible evaluators could slow adoption of the framework without producing comparable information.

Quick briefing

9 min read
  • Anthropic and OpenAI face a Sept.
  • 18 call for independent AI safety audits, with evaluators seeking transparency, access and protection from retaliation.

Institutions Behind the Push

The signatories include Geoffrey Hinton and members of Johns Hopkins University, Stanford University and the nonprofit evaluator METR. Their participation gives the letter representation across academic and specialized evaluation communities, while the AI Evaluator Forum provides the organizing vehicle.

The coalition says only a very small number of groups have the technical credibility, scale and ability to perform this work. That scarcity is an operational constraint, not a measured market share. Until the selection process is disclosed, investors cannot know whether the proposed system would create broad independent capacity or rely on a narrow group of evaluators with limited bandwidth.

Political conditions add another uncertainty. The source reports that President Donald Trump and former AI czar David Sacks have opposed government efforts to regulate AI development, while the letter's authors seek stronger independent oversight. The competing approaches could affect how much responsibility remains with companies, evaluators or public authorities, but the letter does not establish a policy outcome.

Key Debates Investors Should Separate

  • Access versus confidentiality: Evaluators may need sensitive internal data and unreleased systems to assess real risks, while companies must decide how to protect closely guarded technologies. The logistical details of any access program are unknown.
  • Independence versus cooperation: The letter calls for evaluators to operate independently and be shielded from retaliation, yet meaningful inspections require cooperation from the labs being evaluated.
  • Standardization versus flexibility: The A-EF-1 standard is cited as one example, but the coalition says the list of conditions is not comprehensive and needs further standardization, codification and enforcement.
  • External review versus internal controls: The coalition explicitly says embedded evaluations complement rather than replace internal evaluation and broader external oversight.

Companies and Sectors in the Read-Through

  • Anthropic and OpenAI: The two foundation-model companies are the central subjects because the proposal concerns access to their models and development processes. The letter could increase scrutiny of their governance disclosures, but it does not report a financial result or a completed audit.
  • Microsoft: Satya Nadella publicly supported Amodei's evaluator-access proposal. The evidence supports a governance connection, not a quantified revenue, cost or valuation impact.
  • AI software and infrastructure: The broader sector could benefit from clearer independent assessments if enterprise and public-sector buyers place greater weight on verifiable safety practices. That is a conditional mechanism, not an announced commercial outcome.

What to Watch Next

  • Evaluator selection: Any disclosed list of organizations or individuals would show whether the proposed system extends beyond the signatories and how technical capacity is allocated.
  • Access terms: Watch for definitions of “employee-like access,” including company-computer use, employee interviews, sensitive-data review and inspection of unreleased systems.
  • Protection and reporting rules: Investors should look for safeguards against retaliation, publication standards and procedures for handling disagreements between evaluators and labs.
  • Company follow-through: Anthropic, OpenAI and other frontier labs may respond to the letter, but whether they will act is unknown. Any response should be separated from a completed independent evaluation.

Outlook: A Governance Catalyst, Not a Safety Verdict

The bullish case for the AI sector is institutional: independent, transparent evaluations could make safety claims more credible, support more consistent oversight and reduce reliance on self-reporting by a few powerful labs. Public support from Altman, Musk and Nadella gives the access proposal visibility across major technology organizations.

The risk case is execution. The exact signatory count is not disclosed, the evaluator pool may be limited, access boundaries are unsettled and companies may ignore the letter. Even if embedded reviews proceed, the coalition says they cannot address every oversight need. The next meaningful signal is therefore not another endorsement, but a documented program showing who can inspect what, under which protections, and how findings become transparent.

📊 Analysis
Signal  Neutral
Why  The letter raises governance pressure on frontier AI labs, but signatories' exact number, evaluator access and company responses remain unresolved.

This article was independently written by OneDayTrading from public reporting. Read the original (CNBC)

OneDayTrading Editorial Standards

Published by OneDayTrading under its editorial team’s standards. External outlets and institutions named in the article identify reference sources.

Methods, review and corrections
Method
We develop articles and analysis from available public materials, filings and market data, using AI in writing and evidence comparison. Automated checks do not guarantee accuracy. Human review of an individual article is confirmed only when separately indicated.
Analysis basis
We focus on related stocks, sectors, earnings impact, and short-term price catalysts from an investor’s perspective.
Data source
Quotes and foreign/institutional flow data are provided by Korea Investment & Securities (KIS).
Disclaimer
This content is for informational purposes only and is not investment advice or a solicitation to trade.

Bullish or bearish?

One tap to compare your read with other investors.

OneDayTrading Analysis
Editorial signal · key insight
중립

Anthropic and OpenAI face a Sept. 18 call for independent AI safety audits, with evaluators seeking transparency, access and protection from retaliation.

Key theme
AI

OneDayTrading's own editorial assessment. For reference only.

More in AIView all →

© 2026 OneDayTrading. All rights reserved.

US and Korean market news, stock data and analysis for global investors. English coverage combines original reporting with editorially reviewed translations of Korean-market reporting. For informational purposes only — not investment advice or a solicitation to trade any security.