中文版本请见此处
AI 现在不缺品位和直觉,它缺的是持续学习的能力。[AI today isn’t lacking in taste or intuition — what it lacks is the ability to keep learning continuously.] (Liang Wenfeng / DeepSeek Investor Q&A [梁文锋投资者交流会实录])
Guiguzi the ancient Chinese rhetorician, in Book Two of his classic text (English, Guiguzi, Choina's First Treatise on Rhetoric (Hui Wu (trans), 2016)explained that there are situations when, to be in accord with the time one must grasp, or perhaps better, Gauging (摩; mó). Mo is the "probing" or "stroking" chapter of the Guiguzi. The character itself means to rub, polish, or feel by touch — and the technique is exactly that: instead of asking outright what someone wants or believes, you apply a small stimulus (a word, a gesture, a proposal, a show of feeling) and watch what comes back. Internal states, the text argues, can't stay hidden once touched — they surface as a response, the way polishing a mirror brings out its shine, or plucking a string reveals its pitch. The persuader who masters Mo works quietly and indirectly: probe, read the reaction, adjust, probe again, until the other person's true desires, fears, and leverage points are fully mapped — often without them realizing they've revealed anything.
The strategy comes in Chapter 4 of Book II in the classic text, coming right after "Weighing" (Chuai 揣) and directly before "Assessing" (Quan 權). While "Weighing" (Chuai 揣) involves figuring out internal motivations, Gauging (摩; mó) is the active process of testing those calculations through external interaction, mirroring, and verbal probing to prompt a predictable response or action, to be followed by "Assessing" (Quan 權) produces the necessary synthesis of Weighing and Gauging as a function of 道 (Dao--root, path, and sometimes their direction and cognitive conception), materialized as and within the values and premises of the society in which these are deployed. 揣 Chuai (Weighing/Estimating) comes first — a distant, analytical stage. Before any contact, you build a picture of the situation: the other party's power, resources, emotional state, what they love and fear. This is done through observation and inference, largely at arm's length. It's cartography of the mind before you ever set foot in it. 摩 Mo (Gauging) is the contact stage — you take the hypothesis built in Chuai and test it against reality. Where Chuai is theory, Mo is experiment: small probes, calibrated to "their kind" (摩之以其类), reading resonance and resistance in real time and refining the picture until it's confirmed. 權 Quan (Assessing/Weighing) is the deployment stage — now that you actually know who you're dealing with, you calibrate your rhetoric like weights on a scale, choosing different persuasive strategies for the brave, the timid, the greedy, the proud, and so on, in order to tip the outcome in your favor. So the arc is: map (Chuai) → test and confirm (Mo) → calibrate and deploy (Quan). Mo is the hinge — it's what turns a private estimate into verified intelligence you can actually act on rhetorically.
This is precisely the way one might approach the oracularly remarkable products of a long question and answer session given by Liang Wenfeng of DeepSeek in July 2026 (Liang Wenfeng / DeepSeek Investor Q&A [梁文锋投资者交流会实录]). And it becomes more remarkable still as a part of a long conversation, directed outward between the giants of the emerging AI machine system techno-systems: Palantir, Anthropic, OpenAI, Google, Meta and their commentators (for me among the more interesting is Leo Aschenbrenner). Thsi essay considers the product of that Q&A in itself and as part of a broader conversation deeply embedded within the cognitive and phenomenological projections of a number of key actors all projecting their own strategies simultaneously and all in the process producing a surfeit of semiotic contestation that will likely reshape the cognitive framework within the technological mechanisms at the heart of this dialectic revolves.
Introduction:
At a recent investor meeting, Liang Wenfeng, founder of DeepSeek, elaborated on DeepSeek’s organizational culture, open-source philosophy, technology roadmap, and views on the competitive landscape of the AI industry. (Liang Wenfeng / DeepSeek Investor Q&A [梁文锋投资者交流会实录]). The remarks spread widely and a number of sites offered English language transcripts of the remarks some edited to appear in the style of Friedrich Nietzsche's aphorisms (see The China Academy; HTX; WEEX; RootData). I have compiled my own from this site through the former Twitter: HERE. With the assistance of translation apps I have worked through the original text, and include below the original Chinese text and a side by side English-Chinese translation.
The context was a planned DeepSeek IPO. The consequence of the wide distribution of the remarks and its deep or perhaps not so deep crawling through for nuggets by anyone interested in speaking to the remarks, was substantial: "DeepSeek has told prospective investors in its second fundraising round that it’s suspending the deal for now, people familiar with the matter said, days after comments widely attributed to founder Liang Wenfeng about US-Chinese AI competition went viral." (Fortune).
What I add here are presented in five parts: (A) my sense of the four principal insights that these 118 oracular aphorisms afford us; (B) a summary of the aphorisms and oracular statements; (C) Liang Wenfeng's Gauge (摩); (D) The Semiotics of Liang Wenfeng's Guiguzi Strategies; and (E) Liang Wenfeng's Remarks in a Broader Context. In a separate post I will then integrate these oracular aphorisms and those made on 29 July by Mark Zuckerberg for Meta.
A. four principal insights emerge from Liang Wenfeng’s remarks: (1) Strategic Restraint as an Asymmetric Competitive Advantage; (2) Radical Resource Efficiency Can Neutralize Massive Capital/Compute Disparities; (3) A Strict, Disciplined Focus on the "Main Line" to Intelligence; and (4) Non-Traditional, Vision-Driven Culture Over Corporate Bureaucracy.
1. Strategic Restraint as an Asymmetric Competitive Advantage. Traditional tech playbooks prioritize rapid user acquisition, aggressive monetization, proprietary moats (closed-source code), and competitive land-grabs. DeepSeek radically rejects this approach. By maintaining strict restraint—setting API pricing to earn only "reasonable" margins, open-sourcing its frontier models, avoiding "super-app" ambitions, and refusing to view other tech companies as enemies—DeepSeek avoids costly, distracting battles. Counterintuitively, giving away value lowers friction, fosters goodwill, and ultimately maximizes the long-term probability of achieving AGI.
2. Radical Resource Efficiency Can Neutralize Massive Capital/Compute Disparities. The dominant narrative in Silicon Valley suggests that reaching AGI is purely a function of exponential capital expenditure and massive GPU clusters. DeepSeek’s operational reality challenges this "brute force" doctrine. Operating at roughly $1/20\text{th}$ the compute budget of top U.S. labs while remaining only 1 to 2 years behind proves that architectural efficiency, hyper-focused engineering, software optimization (such as writing custom compilers like TileLang to bypass legacy software ecosystems), and lean operations can bridge vast resource divides.
3. A Strict, Disciplined Focus on the "Main Line" to Intelligence. In a market prone to chasing transient hypes (e.g., video generation, 3D asset creation, consumer-facing wrappers), DeepSeek exercises strict intellectual discipline. Liang distinguishes between commercially lucrative distraction and the true pathway to AGI. By treating multimodality as a mere product component and identifying continual learning—rather than sheer scale or world models—as the single vital bottleneck preventing current agents from reaching self-iterating capability, the company concentrates its limited engineering cycles solely on what directly advances core intelligence.
4. Non-Traditional, Vision-Driven Culture Over Corporate Bureaucracy. High-performing AI labs do not necessarily require rigid organizational hierarchies, aggressive KPIs, or intense burnout culture. DeepSeek’s structure relies on a shared, unwritten vision and radical employee trust: granting top researchers up to 50% unscheduled time, avoiding mandatory overtime, relying on consensus-driven leadership, and prioritizing long-term team stability above all else. This relaxed, interest-driven environment is framed as an operational necessity—because genuine scientific breakthroughs in frontier AI require cognitive breathing room rather than top-down pressure.
B. Narrative summary of Liang Wenfeng’s remarks during the four-hour investor Q&A session. These are divided into the following sections; (1) Vision, Organizational Philosophy, and Restraint; (2) Compute Realities, Domestic Chips, and Ecosystem Independence; (3) Commercial Strategy, Open-Source Commitment, and Market Outlook; (4)
1. Vision, Organizational Philosophy, and Restraint. DeepSeek’s core identity is anchored in an unwritten, vision-driven philosophy rather than traditional corporate structures. The company operates without formal KPIs, strict management hierarchies, or performance reviews, prioritizing a mission to achieve General Artificial Intelligence (AGI) for the benefit of humanity over short-term financial returns or aggressive commercial expansion. Liang Wenfeng emphasizes a strategy of deep restraint. Instead of competing with established tech giants to build high-friction "super-apps" or trying to maximize user monetization, DeepSeek deliberately sacrifices immediate commercial land-grabs. This restraint is viewed as a calculated strategy: by remaining focused and avoiding adversarial dynamics across the industry, the company maximizes its long-term probability of actually achieving AGI.
Internally, maintaining team stability is treated as DeepSeek’s single most critical priority. The team views money and physical computing resources as obtainable commodities, but considers cohesive, motivated human talent to be non-negotiable. Recent equity financing provided significant stock options to long-tenured employees, largely mitigating key retention risks. Company decision-making relies heavily on consensus building rather than top-down executive directives. To foster research creativity, management explicitly encourages an unhurried, low-stress work environment. Employees are generally granted up to 50% unscheduled time to explore self-directed research projects without preset deliverables, and overtime is avoided to give researchers the mental bandwidth required for deep exploration.
The AGI Technical Roadmap and Research Focus
DeepSeek views the trajectory toward AGI not as a sudden leap, but as a deliberate, step-by-step technological ladder. The progression moves sequentially from foundational base language models to reasoning paradigms like Chain-of-Thought (CoT), advancing to autonomous Agents, mastering continual learning, reaching a self-iterating technological singularity, and ultimately culminating in embodied physical intelligence. Currently, the primary bottleneck in AI development is that today's agents lack the ability to continuously adapt and learn on the job over extended periods. Unlocking continual learning is viewed as the essential milestone that will allow models to autonomously improve, conduct research, and accelerate subsequent generations of AI development.
Because resources must be strictly prioritized, DeepSeek strictly adheres to this main line toward intelligence. The company deliberately avoids tangential domains like 3D asset creation or video generation, viewing them as lucrative commercial applications that do not fundamentally advance the ceiling of machine intelligence. Similarly, multimodal capabilities are treated as essential product components for end-users rather than core drivers of underlying reasoning ability. To optimize its internal development, DeepSeek prioritizes Coding Agents above all other vertical applications, as strong coding capability creates a compounding feedback loop that drastically accelerates the company's own internal research and model iteration.2. Compute Realities, Domestic Chips, and Ecosystem Independence. Addressing the performance gap between domestic Chinese AI and leading American frontier labs, Liang notes that the divide is driven entirely by resource availability, not a shortage of talent. China and the U.S. draw from essentially the same pool of research talent, but American labs benefit from far greater capital deployment and GPU compute access. Historically, DeepSeek has operated roughly 1 to 2 years behind leading U.S. models while utilizing approximately 1/20th of the compute budget. Moving forward, the goal is to leverage higher computational efficiency to narrow that lag down to 3 to 6 months. Scaling laws remain fully valid, but compute limitations currently restrict how far domestic Chinese entities can scale model size and training data compared to Silicon Valley.
To overcome hardware constraints, DeepSeek has actively pursued ecosystem independence from Nvidia. During the training of DeepSeek V3, the team used Nvidia GPU hardware but entirely bypassed Nvidia’s proprietary CUDA software stack by developing TileLang, a custom high-level compiler that powers their training infrastructure. Looking at domestic hardware, Liang expressed strong optimism for Chinese chip substitution, noting that four Huawei chips currently match the performance of a single top-tier Nvidia GPU. While a 4x hardware gap and a two-year delay in chip manufacturing technology remain, the software ecosystem gap has effectively been closed. Hardware production capacity—rather than software compatibility or adapter ecosystem barriers—remains the sole operational bottleneck for domestic compute.3. Commercial Strategy, Open-Source Commitment, and Market Outlook. DeepSeek approaches commercialization with a philosophy of cost-plus pricing rather than margin maximization. API pricing is calibrated to cover hardware operational costs and recover capital expenditures within approximately ten months, intentionally offering prices far below what inelastic market demand would allow. Enterprise (To B) revenue is expected to reach hundreds of millions of dollars, which, paired with a growing consumer user base, puts the company on a fast track toward net profitability. Even in a hypothetical worst-case scenario where technological progress plateaus, selling API access alone would be sufficient to sustain a profitable, publicly traded enterprise.
Furthermore, DeepSeek remains firmly committed to an open-source strategy, intending to release even its most powerful frontier models to the public. Liang argues that closed-sourcing provides no inherent competitive moat, as model deployment, cost optimization, and operational efficiency present formidable barriers to entry even when weights are fully shared. Crucially, the models DeepSeek deploys internally for its own API services are identical to the weights released to the open-source community. On the global competitive stage, Liang envisions an industry where no single player holds a monopoly or extracts windfall profits. Instead, intense competition will turn cost efficiency and execution speed into the primary market differentiators, with Chinese companies positioned to deliver global AI capabilities at significantly lower price points.
What makes the summary odd is its usefulness. It seeks to make sense of a set of flowing aphorisms that one can group and regroup as one likes to mold the aphorisms and oracular statements into something that maybe it is not--like the summary I offered above. It serves a purpose but exposes another--the desire and mechanics of seeking to impose meaning on something that means what it says and follows its own discursive rhythms which may or may not have meaning beyond the aphorism itself. Yet there it is, a necessary hallucination in aid of meaning that must be imposed. 大模型的幻觉问题比较影响用户的体验。幻觉问题也是有一个方法可以解决的,但是这是一个长命题。幻觉问题可以认为是一个可以通过更好的 Post-training 解决的,是一个能解、能够改善的问题。 [ The hallucination problem in large models significantly affects user experience. There is a way to address the hallucination problem, but it’s a long-term challenge. Hallucination can be seen as a problem that can be addressed and improved through better post-training. ] (Liang Wenfeng / DeepSeek Investor Q&A [梁文锋投资者交流会实录])
C. Liang Wenfeng's Gauge (摩)
Liang Wenfeng's transcript is not just information about strategy — it's itself an act of strategic communication, and the framework can be pointed at it two ways.
1. The document as an act of Mo directed outward. Notice the setup: a company that explicitly refused to speak — "never raise outside funding, never go public" — breaks four years of near-silence in a single four-hour session, immediately after closing a RMB 50 billion raise. That timing is not incidental. A silent company is illegible; markets, competitors, and the state all have to guess at DeepSeek's intentions. This session is a controlled probe in the other direction — Liang is the one being "gauged" by 118 targeted questions, and his answers are calibrated releases of information timed to a specific moment (post-raise, pre-scale-up) when reassurance to a very specific audience — his new investors — has maximum value. It reads like Quan already at work: he isn't giving a uniform message, he's weighting different reassurances for different anxieties in the room — team retention for people worried about talent flight, chip strategy for people worried about export controls, restraint-as-strategy for people worried he'll monetize recklessly and alienate the ecosystem he depends on for goodwill.
2. The document as raw material for our own Chuai–Mo–Quan. If you're the reader trying to actually assess DeepSeek rather than just absorb the narrative, Guiguzi suggests treating his stated positions less as facts and more as probe responses to be weighed. A few examples of where his own language, tested against itself, reveals more than any single line: He repeats "restraint" (克制) as almost a mantra — "Restraint is itself a strategy. Sometimes you can give something up in exchange for more of something else." Guiguzi's Chuai stage would ask: what is restraint actually buying him? He answers this almost directly later — restraint is explicitly framed as a probability-maximizing move toward AGI, not altruism for its own sake: "what I prioritize is how to increase the probability that we succeed." The "goodwill" framing and the cold optimization framing sit side by side without friction for him — that's worth noting rather than resolving.
On team stability he's unusually blunt that money is a solved problem and the only remaining risk is retention — "Our single greatest core interest is maintaining the stability of the team... As long as I can keep the team stable, I will definitely succeed." That's a rare moment where a Chuai-style estimate (what does this man actually fear?) gets a direct, unguarded answer instead of a rehearsed one — arguably a slip induced by the interview format itself, exactly the kind of unplanned resonance Mo tries to elicit.
On competitors he consistently downgrades rivals' advantages as temporary — Anthropic's lead over OpenAI, OpenAI's near-term dominance, even Nvidia's CUDA moat — all described as eroding or "just a phase." A Quan-style reading would flag this as a consistent rhetorical pattern (not a one-off claim) aimed at reassuring investors that no competitor's current position is fixed, which is precisely the reassurance an investor writing a check at a 367.5B RMB valuation, after two years of "we'll never raise," most needs to hear.
The broader point the Guiguzi framework surfaces: this transcript shouldn't be read as a transparent window into DeepSeek's strategy so much as a document produced by someone highly practiced in exactly the technique described in 摩 — reading a room and calibrating disclosure to it. The most interesting analytical move isn't taking the content at face value, but asking what stimulus (the funding round, the export-control pressure, the domestic-chip narrative) each answer is a response to, and what that implies about what he still hasn't said.
D. The Semiotics of Liang Wenfeng's Guiguzi Strategies.
Semiotics sharpens what's actually happening in each of these three moves, because the whole Guiguzi method rests on treating a person as a sign-system rather than a transparent container of intentions. The core semiotic assumption underneath all three chapters is that internal states (情, feeling/disposition) are not directly accessible — they only become knowable through externalized signs: words, tone, posture, timing, silence. This is already a semiotic move in the Peircean sense: the internal referent is never given directly, only inferred through indices (involuntary, causally-linked signs — a hesitation, a flush, a repeated word) and symbols (conventional, coded signs — the actual vocabulary chosen). Guiguzi's whole method is a discipline of reading indices through symbols, on the assumption that no one fully controls both channels at once.
Chuai (揣), semiotically, is code-construction before contact. Before you can interpret any sign a person gives off, you need an interpretive frame — a working model of what kind of person you're dealing with, what their signs are likely to mean. This is structurally close to what a Saussurean would call establishing the paradigmatic field, or what a hermeneuticist would call pre-understanding: you can't decode without a code, and Chuai is the stage where that code gets built, entirely from a distance, through observation and inference rather than direct exchange.From this the sequence--Guigizi's sub-system block chain can be read this way: build a code (Chuai) → test it by manufacturing signs (Mo) → produce new signs calibrated to that code (Quan) — a full semiotic loop, not just three stages of "getting to know someone." Applied back to the transcript, this adds another layer to the analysis:
Mo (摩) is where the sign gets manufactured, not just observed. This is the crucial semiotic move that distinguishes it from mere reading. Natural signs are often absent, suppressed, or ambiguous — so the practitioner doesn't wait for a sign, they induce one: inject a stimulus, calibrated to the target's "kind" (摩之以其类, i.e., in a code the target will actually respond to), and treat the resulting reaction as data. It's closer to Peircean abduction than passive observation — form a hypothesis in Chuai, then perturb the system in Mo to force it to emit a legible sign that confirms or corrects the hypothesis. The requirement that the probe be pitched "to their kind" is itself a semiotic constraint: sender and receiver need a shared code, or the elicited sign comes back as noise.
Quan (權) is re-encoding for effect. Once you have a validated read of the other party, you stop decoding and start producing — selecting rhetorical categories (appeals to fear, ambition, loyalty, pride) the way you'd select weights on a scale, each keyed to trigger a specific decoding on the other end. In modern terms this is the encoding half of Stuart Hall's encode/decode model: Chuai and Mo are the decoding operations performed on the other person; Quan is the encoding operation performed for them, using exactly what decoding revealed about how they process signs.
愿景 (vision) functions almost as a floating signifier across the document — repeated relentlessly, explicitly said to be unwritten: "This vision isn't even written down anywhere... it never has been." Semiotically that's notable: a symbol with no fixed signified is maximally efficient for consensus-building, because every listener — employee, investor, state regulator — can attach their own referent to it without contradiction. It's a sign selected in Quan precisely for its interpretive elasticity, not despite it.
Numeric precision as a sign of certainty. "Four Huawei chips equal one Nvidia chip" and "twelve to eighteen months, or six to twelve months" behave less like raw data than like rhetorical icons — precision itself is the signal, standing in for confidence and control, regardless of how the number was derived. Worth flagging as encoded output (Quan) rather than transparent index of fact.
The curatorial layer. The transcript states outright that it's "118 remarks, organized by theme, retaining only the substance of what he said" — meaning we're not even reading Mo's raw elicited signs firsthand; we're reading a second-order recoding, filtered by whoever assembled the document, on top of Liang's own first-order recoding of what four hours of Q&A produced. Any semiotic reading of this text has to hold that double-encoding in view — the "signal" we're decoding has already passed through two encoders before it reached us.
Here we move from Guiguzi as rhetoric, through semiotic dialectics within that rhetorical cage, to the computational expression of that dialectic that then mirrors the machine system that is the object of the exchange. In its essence one arrives, yet again, at inductive systems emerging from iterative mimetics that stars as flat sequential nodal movements in one direction and then acquires a layered polyphonic (in regulatory cognitive spaces polycentric) element that produces the end product sought by seeking to sketch out with 118 oracular nodes the path toward its realization.
E. Liang Wenfeng's Remarks in a Broader Context.
1. DeepSeek Within an AI Peer Group Conversation. I have not considered Liang Wenfeng's remarks/aphorisms/oracular pronouncements in a vacuum. And, indeed, by the time they were made Liang Wenfeng had had several months (it is a closely knit community worldwide) to digest the oracular pronouncements and interventions of his peers in the United States, peers who also found it hard to keep their thoughts to themselves (see, Palantir, OpenAI, Anthropic, and Leo Aschenbrenner). I considered these in recent lectures at East China University of Politics and Law (Lecture
7— AI Narratives and the Future of AI-Human Regulatory Structures from a
Human, Machine Computational, and Machine Quantum Perspective;
Palantir; Anthrop/c; OpenAI--for the Lecture Series: AI Governance in
Comparative Perspective, Theory and Practice: China, U.S. and E.U.; Lecture Series Homepage HERE).
My Lecture 7 shifted the AI governance discussion from state regulators to private-sector actors, treating public statements from Palantir, Anthropic, OpenAI, and independent essayist Leopold Aschenbrenner as competing "oracles" about who should hold authority over AI's future, rather than as technical policy papers. The lecture's governing premise is dialectical: AI systems are produced by the political orders that build them, but recursively reshape those same orders' institutions, cognitive habits, and norms. In that Lecture, as in the analysis of the underlying texts, I framed the four texts through the allegory of Sophocles/Cocteau's Oedipus Rex — Oedipus as confident problem-solver, Creon as administrative ruling class, Tiresias as a technical intelligentsia serving power rather than truth, and Jocasta as the dissenting voice who exposes the oracle's lie.
The four narratives are read as four different "governance objects" for the same technology: Palantir treats AI as an instrument for reconstructing the state from within — the state must be reorganized around AI-enabled visibility and coordination, with a Silicon Valley engineering elite as a legitimating vanguard (what I have called "techno-Leninism"). Anthropic externalizes AI into a civilizational, US-China contest over compute, export controls, and model "distillation," treating AI capability itself as contested territory rather than as an agent, with 2028 as the decisive horizon. OpenAI proposes "transformative preservation" — deep societal change managed through public-private partnership so that legitimating institutions appear undisturbed even as their substance changes. Aschenbrenner (Situational Awareness) radicalizes all three by treating a national-security state as near-inevitable once recursive self-improvement produces an "intelligence explosion," leaving only the question of whether humans or autonomous systems ultimately direct it.
Reread computationally and quantum-computationally the same four texts suggest a convergent finding across all three passes: none of the four architectures disputes that human authority should be nominally preserved, but all four converge on structures in which human authority becomes an "interface property" — a legible, answerable-to layer — while operative agency migrates elsewhere (an administrative elite, contested infrastructure, a technical minority, or the system itself). The quantum pass adds that human governance's sequential, nodal, and irreversible temporal structure is structurally incommensurable with computational time, so governance corrections systematically lag a self-accelerating capability trajectory.
Where might Liang Wenfeng fit into that conversation of peers? Liang Wenfeng's remarks sit outside of the four-narrative American taxonomy, but they engage several of the same structural questions — authority, timing, harm/risk, and legitimacy — from a markedly different institutional and geopolitical position.
On authority and governance structure. Where Palantir locates authority in a reformed state apparatus fused with an engineering vanguard, and Anthropic and OpenAI locate it in state or public-private coordination, Liang locates DeepSeek's internal authority in consensus rather than command: he states the company has no KPIs, is "vision-driven," and that his own influence "is built on the foundation of consensus". This is nearly the inverse of Palantir's proposition that judgment and hierarchical discrimination among values must be restored to a ruling elite; Liang instead describes a flat, half-unscheduled research culture explicitly organized to avoid administrative control. Notably, Liang's account never assigns China's state a governing role in DeepSeek's mission — the company's felt obligation is commercial survival ("the government won't give us a single cent") rather than state-directed purpose, a contrast with my account of Anthropic's document, which frames AI governance as inseparable from a state-versus-state contest for normative dominance.
On the US-China frame specifically. Anthropic's narrative casts China as a strategic adversary whose gains stem from talent, loophole exploitation, and "distillation" of American models, with 2028 as a resolving horizon after which either democratic or authoritarian norms will govern AI globally. Liang's remarks invert the vantage point of that same contest: he describes DeepSeek as roughly one to two years behind the US while using only about a twentieth of US compute, attributes the entire capability gap to compute and capital rather than talent, and forecasts that China's comparative advantage will be cost and production scale rather than a different kind of intelligence. Where Anthropic frames the contest as a fight over whose values set global norms, Liang frames it in market terms — China will make AI "the cheapest," with pricing that only earns "a reasonable return" rather than maximum extraction — a rhetorical register closer to industrial competition than the securitized "civilizational competition" vocabulary I have attributed to Anthropic.
On timing and the "singularity." Liang's AGI roadmap — chain-of-thought, then agents, then continual learning, then a self-iterating "singularity," then embodied intelligence — is structurally similar to Aschenbrenner's recursive self-improvement/"intelligence explosion" logic, in which AI automating AI research compresses years of progress into a shorter span. But Liang explicitly resists Aschenbrenner's explosive framing: he insists the "singularity" is "not really a singularity" but "a gradual process... not a sudden leap," even while conceding it is habitually described in dramatic terms. This directly contradicts Aschenbrenner's discrete compressed 2027–2028 event horizon, and, to a lesser degree, to Anthropic's. Within the discursive framework I developed for Lecture 7, then, Liang's timing model resembles the "continuous, ex ante" administrative correction I have attributed to Palantir and OpenAI rather than the discrete terminal-horizon models of Anthropic and Aschenbrenner.
On legitimacy and risk framing. The American narrative's legitimacy warrant circles around key conceptual organizing concepts: patriotic moral debt (Palantir), defense of democratic process (Anthropic), broad-based shared prosperity (OpenAI), and sheer survival (Aschenbrenner). Liang's legitimacy claim is closer to OpenAI's "shared prosperity" register but grounded in restraint rather than democratic participation: he repeatedly frames deliberately not maximizing DeepSeek's share of AI's payoff — through low API pricing, continued open-sourcing of even its strongest models, and explicit willingness to help competitors such as Alibaba, Zhipu, and Moonshot AI — as the strategy most likely to increase the probability of reaching AGI at all. This stands in sharp contrast to Anthropic's zero-sum "distillation" framing, in which a rival's extraction of capability from a leading model's outputs is described as adversarial capture; Liang treats the analogous risk — competitors freely deploying and even improving on DeepSeek's open-sourced weights — not merely as tolerable but as a documented policy goal irrespective of the classical computational reading's point that it entangles DeepSeek's fate with the actors who redeploy its models.
On the authority/agency decoupling. A key structural finding suggests that all four American texts formally retain human authority while operative agency migrates to an administrative elite, contested infrastructure, or the system itself. Liang's account offers a partial counter-case worth flagging rather than a clean rebuttal: he explicitly ties DeepSeek's continued viability to "keeping the team stable" as its "only core interest," identifying a small set of senior researchers (whom he says make up roughly half the company, concentrated on data annotation) as the load-bearing agents of the enterprise. That is structurally similar to what my computational reading calls Palantir's "entangled subsystem" of an engineering vanguard whose own state cannot be specified independently of the system it supervises — except that Liang frames this concentration as a talent-retention and morale problem solved by financing-round equity, not as an emergent administrative authority displacing collective human judgment. Whether Liang's account of consensus-based, KPI-free governance would itself survive what might be called my "decoherence critique" (i.e., whether "vision-driven" consensus is a durable control property or, per the quantum reading, merely an "interface property" masking concentrated operative agency in DeepSeek's core research team) is a question Liang's remarks do not directly address — the document offers no equivalent second-order reflection on whether its own account of internal governance could itself be characterized as a legitimating narrative rather than a description of operative control.
Net Comparison. Within these American AI firms' governance narratives, each firm's stated commitment to preserving human authority lies a structural tendency to relocate operative control elsewhere. Liang's remarks are a first-order narrative themselves — not a policy document self-consciously arguing for a governance architecture, but an internal account of DeepSeek's strategy, culture, and market position. Read through my own analytic lenses, the DeepSeek document would likely occupy a position distinct from all four of the American texts I consider: it neither embeds AI within state administrative reform (Palantir), nor casts AI capability as contested geopolitical territory to be defended (Anthropic), nor proposes a formal public-private error-correction architecture (OpenAI), nor forecasts an inevitable security-state capture (Aschenbrenner). Instead, Liang frames restraint, open-sourcing, and consensus governance as strategy for a private, commercially-exposed firm operating from a position of resource scarcity relative to the US — a register closer to entrepreneurial pragmatism than to any of the four oracular postures I catalogue, even though it shares with Aschenbrenner a recursive self-improvement roadmap and with Anthropic an explicit US-China compute framing.
2. DeepSeek's Political-Cognitive Platform: Operating Inside the Guided State. In my Lecture series (Lecture Series Homepage HERE) I describe the Chinese regulatory environment "The Guided State", and Lecture 7 itself, in glossing Anthropic's narrative, contrasts the American framework — "organized around markets and national security" — with "a Chinese framework organized around what the document terms 'Socialist Modernization' driven by state-directed, high-quality production," and situates the American liberal-democratic project as defending its "lebenswelt" against "the imaginaries of Marxist-Leninist successor states". This is the essential structural fact that has to frame any reading of Liang Wenfeng's remarks: DeepSeek does not operate within a system where market autonomy, decentralized private ordering, and firm-level self-direction are constitutionally protected defaults. It operates within a system whose foundational premise is Party leadership over the economy, in which market mechanisms are instrumentally tolerated and steered rather than treated as an autonomous sphere prior to or independent of collective political direction. Enterprises of DeepSeek's scale and strategic significance in China typically function under and alongside embedded Party organizational structures, and the Party-state has, in recent years, repeatedly demonstrated both the capacity and the willingness to discipline private technology firms perceived as accumulating disorderly, unaccountable economic or social power — the treatment of Jack Ma and Ant Group after 2020, the restructuring of Didi, and the broader "common prosperity" campaign against the "disorderly expansion of capital" are widely documented, publicly known instances of this dynamic. This is general background context rather than something drawn from the uploaded materials, but it is necessary to read Liang's remarks accurately, since his rhetoric is being spoken into precisely that environment.
Against that backdrop, several features of Liang's remarks read very differently than they would if spoken by a Silicon Valley founder. Liang repeatedly and emphatically disclaims the pursuit of dominance, scale-for-its-own-sake, and market power. He states that DeepSeek never intended to "become the next ByteDance, the next Tencent" and has no wish to compete with "any internet giant or small company". He describes the firm as having "no organization" in the ordinary sense, governed by "vision" rather than "KPIs," with no performance review at all. He states that "our single greatest core interest is maintaining the stability of the team" — "you could even say it's our only core interest". He frames open-sourcing the firm's strongest models, helping competitors including Alibaba, Zhipu, and Moonshot AI "do better," and pricing API access at "a reasonable profit" rather than a profit-maximizing rate, all as deliberate, principled choices rather than commercial necessities. And he generalizes this into an explicit political-economic claim: "whoever takes more will be beaten by whoever takes less" — that a vision oriented toward capturing more market share or profit is itself a competitive liability.
Read acontextually, this could be mistaken for standard Silicon Valley founder mythology. Read against the guided-state backdrop, it performs a much more specific and consequential kind of signification. Liang's language of "restraint" (克制) is explicitly theorized by him as strategic rather than merely temperamental: "restraint is itself a strategy. Sometimes you can give something up in exchange for more of something else". What DeepSeek gives up, on this account, is overt scale, dominance, and profit-maximization; what it purchases is something Liang never states outright but that the guided-state context makes legible — continued latitude to pursue an extraordinarily ambitious, resource-intensive, and politically sensitive project (building AGI) without triggering the pattern of Party scrutiny and disciplining that has met other Chinese technology firms perceived as amassing autonomous, unaccountable power. His insistence that DeepSeek is "vision-driven" rather than rule-driven, and that its authority rests on "consensus" rather than unilateral command, performs non-threat in a system whose default posture toward concentrated private authority is suspicion. His statement that the firm's vision "isn't even written down anywhere" and "lives in the way we do things" is, in a Western frame, a claim about organizational culture; in the guided-state frame, it is also, functionally, a claim about the absence of any documented alternative locus of authority that could be read as rivaling or displacing Party-sanctioned direction.
This is the sense in which the DeepSeek text should be read as a distinct fifth governance object in my Lecture 7 typology, produced by, and legible only against, a political order fundamentally different from the American "markets state" context that produced Palantir, Anthropic, OpenAI, and Aschenbrenner. Where the American texts are oracles addressed to a public and a state apparatus that must be persuaded to accept unusual concentrations of authority, Liang's remarks are an oracle addressed to a Party-state apparatus that must be persuaded that no unusual concentration of authority is occurring at all.
(text produced in collaboration with Claude and Harvey AI)
融资 500 亿后首次开腔:梁文锋 4 小时讲透 DeepSeek(118 条实录)
说过「不融资、不上市」的 DeepSeek,这次融了超 500 亿元,投前估值约 3675 亿元。围绕这次融资,梁文锋在一场近 4 小时的投资者交流会上回应了一切:愿景、AGI 路线图、算力、国产芯片、竞争格局、开源,共 11 个话题。以下是他的 118 条发言,按主题整理,只保留讲话内容。
01 愿景与克制
我们一开始来做这个公司,初衷没有想到说我最后要赚多少钱,要到资本市场上去,要上市,要怎么样的。最开始的几十个人完全没有这么想过,如果他这么想,他就不会来。
我们是怀着一个对这个世界非常大的善意来做这个事情,我们觉得这是对人类有用的,这是一个金钱以外的事情。我们出发的初衷、我们的愿景,以及我们保持到现在的这个愿景,不是按照一个商业利益最大化的方式来做的。
管理一个大公司,靠的不是你的规章制度,靠的是愿景。愿景不是挂在墙上的标语,愿景是你怎么做,不是怎么说,就是你怎么实际运行。
我们是没有组织的,就是愿景驱动的,靠一个愿景来组织。我们并不是以一个"我要实现什么 KPI、没有考核"的方式来做,只有愿景。
这个愿景甚至也不是成文的,并不是写出来的,没有写出来过任何东西。这个愿景是在我们做事情的方法、我们对待这个世界的态度里。
我们并没有非常多的其他优势,我们没有什么本事,我们并没有比别人有钱,也没有说我们人员比其他公司更好,其实没有的。我们两年前成立这个公司的时候,我们又没有很多钱,又没有很多卡,又没有什么知名度,又没有什么号召力,我们就是一群非常平凡的人。
你越克制,可能就越容易做成,或者说至少到目前为止是印证的,到目前为止是能解释得通的。否则没有办法能够解释为什么我们能够做成:我们并没有什么武器,起点又非常低,资源又非常少,我们的人其实也就是随机的一群平凡的人。
AI 这个事情太大了,利益太大了。我们非常克制,只要能够做成,最后利益都会非常大。你随便分一点,利益就非常大,所以现在根本不用考虑拿这里面的哪一部分利益、怎么拿,我觉得根本不用考虑这件事情,因为这个利益足够大了。
我们去年春节用户突然很多,但是我们并没有去追求我要留这些用户,或者说拿这些用户来变现,或者说我要去抢这些商业利益,在用户上面兑现。我们没有去抢用户,没有去赚钱,但我们很努力想办法把用户服务好。
我们并不会有这样的想法,说我要做成下一个超级 App,然后我要去跟谁竞争,我要做成下一个字节、做成下一个腾讯,完全没有这样的想法。我觉得后面的 AGI 机会应该是非常大的,后面的 AGI 机会永远是非常大的。
克制是一种战略。就在于有时候你可以舍弃一些,来换更多其他的东西。不开源这个事,其实也是一样的,也可以认为是我们的压力,也可以认为是我们的让利。
这种克制,我理解是这种克制从长远上来讲,能够增加我们做成 AGI 的概率。在考虑一件事的时候,我是毫不怀疑 AGI 会有非常大的商业价值。那么在这个基础上,我优先考虑的不是我怎么多加一点份额,我怎么多拿一些份额,我优先考虑的是我怎么增加我能够做成的概率。
我们一直非常克制,不愿意跟任何一家互联网大厂或者小厂成为对手。我希望我能够给他赋能,或者希望我能够协助大家去做这个事情,希望能够帮助大家做这个事情。
我觉得我们之前秉持这种态度,其实我们并没有因此而少拿到任何东西,并没有因为我开源,并没有因为我们的善意或者说我对其他人提供帮助,而导致我们少拿了任何东西。反而可能还有加分。这个看起来违反直觉,但它确实是这样的。
我们是以 AGI 为目标,但是我们一直在做商业化,所以我们才有 C 端的用户,才有 B 端的收入。从历史经验来看,这个策略是成功的。
02 AGI 路线图
如果说你能够把一个问题描述得很清楚,给它完整的上下文和指令,它已经超过人类了。但这里有个定义,有个前提是:你给它完整的上下文,给它完整的指令。
AI 并不能够替代你的员工。但如果说 AI 具有持续学习的能力,它跟你的员工一样,到公司学习两个月,那么就可以替代天下的人了,所以我们离下一步还差一个持续学习。
AI 的发展,我们可以理解成它是一个阶梯。去年走的阶梯是思维链。因为我们发现,通过思维链的方式,可以让智能达到一个更高的水平。
今年的阶梯就是 Agent,因为我们发现,用 Agent 的方式,即便多些事情也可以做,它的能力范围会更大,它的智能上限会更高。Agent 要用到 CoT,然后 CoT 也要用到前面的阶梯,前面的阶梯就是语言模型,所以它并没有一步是白走的。
在 Agent 之后,我们觉得应该要解决的问题是持续学习,就是怎么让模型可以持续地学习,而不是说你要给它一个很强的训练,它应该能够像人一样做一个比较长时间的持续学习。
持续学习之后,可能我们就会来到一个奇点。这个奇点就是,当这个模型能够持续学习之后,它已经能够做人类能做的所有事情了。它就能够自己开发自己的版本,能够自己再研究,然后开发自己的下一个版本,开发更好的人工智能模型。
这个奇点,它并不是一个奇点,它也是一个渐进的过程。这个过程可能也是一个比较长的渐变,它不是个突变。但是习惯性地,我们都认为它可能是个奇点。
这是我们的推测,这是我们觉得这个时间表应该是:先解决学习的问题,然后再到那个智能的奇点,能自我迭代的奇点,然后才是具身智能。到具身智能之后,它就走进现实世界,可以给你做家务,可以给你养老。
如果说我们先解决持续学习,再解决那个自我迭代的奇点,再解决具身智能,这个路上就很轻松。因为到后面之后,你可以用前面的技术来帮助开发后面的技术。
我们只做 AGI 的主线。AI 领域很广泛,有很多东西我们觉得它不在这个主线上面,比如说 3D、视频生成,我觉得可能跟智能的主线没有太大的关系,我们不会去做。
视频生成一开始出来的时候就很火,好像这是必须要做的,如果你不做,你就不是一个 AI 公司一样。所以我就很奇怪,这个其实你只要仔细去想一下,它跟智能的路线图是没有什么关系的。
在商业上,它是个好生意,在商业上是个好生意。但是这跟智能没有什么关系。我们不会因为它是一个好商业而去做它,我们只会因为它是智能路线图上的东西,才会去做。
从我们的判断,世界模型和智能还不是现在这个阶段最重要的事情。最重要的才是 AI 训练,以及 AI 训练之后怎么解决持续学习。这是我们公司的判断,当然每个公司它的判断是不一样的。
我们现在比较相信一个叙事是,AI 可以加速 AI 的研究。就是说,它不是线性的,因为你可以用 AI 来加速你自己的研究,所以它到后面可能是非线性的。
我觉得具身肯定还是要进入,最终具身。因为对一个正常人来讲,他的需求并不是电脑,对不对?因为正常的人,他吃喝玩乐、衣食住行,他不需要电脑。他需要的是,所以他还是需要具身智能来解决具体人类的需求。
我们希望 AGI 能做什么呢?它能帮我迭代下一版模型,能帮我迭代下一版模型一样。如果有了具身之后,我们希望它做的也是,让它来迭代下一版的具身,它来做下一版机器人。
下一代模型的核心能力,必须要有持续学习的能力,它才能叫下一代模型。在那之前,我们能做的就是降低成本,然后效果做得更好,速度做得更快。但是要有大突破,它应该是具备持续学习的。
现在的 Agent 的能力受限,是因为它不能持续学习,它不能有效地持续学习。如果说能够先把持续学习做完,那 AI 的能力是非常强的,它能够非常大地提升我们自己研究的效率。
持续学习先做出来,通用智能可能就很容易了,用它来做就很容易。所以我说这是一个我们比较希望看到的结果,我们比较省力,我们就轻松。否则现在你要去人工做通用智能,它是一个比较累、比较苦,是一个数据密集、人力密集的事情,性价比也不高。
03 团队与人才
我们前面的经历给我的启示是,AGI 这个愿景是很强大的。这个人才优势不是说我的人比他更聪明,而是这些人才我怎么组织起来,怎么激励他,然后怎么合作。
你把聪明的人聚在一起,并不是说他自然而然就能够合作,自然而然就能够非常有激情地去奔一个目标、去完成的,所以你需要一个愿景。
我们最大的核心利益是要保持团队的稳定性。这是我们最大的核心利益,甚至可以认为是唯一的核心利益。只要我能够保持团队的稳定性,我一定能做成,一定能做成 AGI,就这么简单。
钱肯定不是问题,资源不是问题,其他要素都是容易获得的。对我们来讲,只有一个核心利益,只有一个没法退让的:我们必须要保持团队的稳定性。
这也是我们面临的一个非常大的挑战,或者说,我觉得是最大的风险。当然,这个风险随着我们近期的这一次融资,得到了比较大的解除。因为大家拿到的期权都还是比较多的,金额还是比较大的。
从团队稳定性来讲,只要最重要的一些员工、最老的员工能够稳定,那么其他人是不太会走的。其他人哪怕期权少一点、收入少一点,他也不会走。因为他不是全奔着钱来的,大家都希望在一个能够做成 AGI 的环境里面去做这个事情。
其他都是时间问题,其他最多导致我们晚半年、晚一年,但是不会说做不出来。肯定是不缺钱,肯定是不缺资源,其实这些都是不缺的。
我们跟美国的差距主要在资源上面,然后人上面差距不是很大的。人上面几乎没有差距,因为就是同一批人,可能是中国人。中国人出去的时候,有一些人留在国内,有些人留在国外,有些人去国外,他并没有说是聪明的人去国外,没有的。
人才不是瓶颈,资源是最大的瓶颈。资源首先影响到人才培养,因为算力少,我们能做的实验机会比较少,所以我们的人才整体上比美国有差距。人才的差距,本质上也是因为算力的差距。
AI 人才的短缺也是阶段性的,并且我们已经看到,大幅度被缓解了。因为 AI 人真的不缺,每个公司很快会把人培养出来,培养人是很快的。
国内现在做模型的公司有点太多了,还是太多了。美国可能就三家,中国做基模的东西太多了。最终一定是不需要那么多人去做基模的,一定会收敛。
我们公司的管理其实是两条线:一条线是从上到下,一条是从下到上。从下而上,就是每个人自己想做什么,自己做,没有人管他,没有 KPI。
一般我们希望员工还有一半的时间,是不被安排的,他想做什么就做什么。这是一个研究的范围,让他可以自己去探索,按照他觉得什么重要,他去探索什么,没有前置的要求。
我们一般也不太加班。加班有两个原因。第一个是,做研究是需要一个比较松弛的环境。你如果逼得很紧,就没法做研究。因为既然就是要你自己有这个兴趣,你自己平时要去想这些问题,所以得是在一个比较松弛的环境里,才有可能能够探索。
第二个是,我们非常聚焦。我们非常聚焦,就意味着我们要做的事情很少。那我就没那么多事情要做,我就不需要加班。这个跟前面的克制是一脉相承的。
我们公司整体上是建立在共识的基础上的,我并不是说我一个人决定所有事情,而是我要寻求共识。我在公司内的权威以及在公司内的影响力,是建立在共识的基础上的。
这个决策机制其实是一种寻求共识的机制,这并不是说我能够推动一个什么事情,它一定是共识,我才能够推得下去,然后我才会去推。
随着人员的增加,我们会做这个调整。应该是我们马上就得做这个调整,因为我已经在做这个调整。已经不做这个调整的话,很多事情是没法推进下去的。确实有很多部门,是应该有组织架构的。
04 算力与资源
我们需要多少卡?现在肯定是越多越好。在我们能够承受的范围内,肯定是卡越多越好,这是毫无疑问的。所以我们现在策略是,在合理的价格里面,能买到多少卡就买多少卡。
实际上,要把这么多钱花完非常不容易,买不到那么多卡,很难买,而且价格也很高,也不能说花非常高的价格去买,还得确保这个价格是合理的。如果今年能够花掉两百亿,那么就属于我们的采购部门业绩超级好了。
我们跟美国之间最大的差距是在资源上面。算力资源一方面是国内本身卡就买不到,另外一方面是我们的资本投入比美国要少。我们在资本投入层面上少了很多,基于人才的工资在这里面占的比重是很低的。你看他们开的薪水一个亿美金这种,但算起来,人才的薪水还是占比很少,大头还是算力。
我们看到的所有区别,包括人才的区别、模型能力的区别、应用的区别,可以认为都是因为算力资源上的区别。
我们跟美国的差距可能是落后美国 12 个月,落后美国可能 12 到 18 个月,或者说 6 到 12 个月。反正简单说,就是落后美国两年,然后只用美国二十分之一的算力把这个事情做出来。
这个叙事就是落后一到两年,但是只用它二十分之一的算力。那么未来我们要把这个叙事改写,就是我们用它几分之一的算力,但是把这个时间缩得更短,缩到 6 个月、3 个月,我觉得这是一个目标。
Scaling,我们是信 Scaling,肯定是规模越大,效果越好,能够解锁更多的功能。阻止我们 Scaling 的其实就是算力,并不是我们不想 Scaling,是我们没有那么多的算力去做这个 Scaling。
我们训练这么大的模型,并不是因为我觉得这么大的模型就够了,而是我刚好有这么多资源。我是按照我的资源来算,我能够接受、能够训练的模型是多大,是这样算出来的,并不是这个模型就够了。
硅谷在说 Scaling 到头的时候,那是对硅谷来说;对中国人来讲,我们离那个还很远,我们根本就没有 Scaling 到那个程度。这个 Scaling 包括数据的 Scaling、模型规模的 Scaling,然后训练成本。
05 国产芯片与生态
英伟达 CUDA 的护城河在快速地被瓦解。一方面是现在有了 AI,然后有了 AI 之后,我要建立起这个生态比以前容易很多了,因为 AI 可以写代码。
现在计算卡的市场已经比游戏卡更大了,就没有理由这两个还需要耦合起来。现在的趋势是,以后就不再耦合了。那么专用芯片,不管是华为还是英伟达自己,以后都是专用芯片,都不是之前的这些东西了。
国产 AI 芯片替代现在是有个历史性机会的。我们认为,未来一年之内,我们能够看到有一个事情被验证:国产芯片的生态完全没有问题。之前认为是有问题的,认为是用不起来、不好用,但是未来我觉得一年之内,我们能够扭转这个认知,或者会用事实来扭转这些。
国产 AI 芯片的硬件和生态都没有问题,唯一有问题的是产能不够。国产卡适配这一点,没有障碍,英伟达挡不住。如果是在一个正常的商业环境里面,我能买到英伟达的卡,那么国产替代是比较难的;但是在英伟达的卡买不到的情况下,所有人都迫不得已,都要去做国产芯片。
V3 训练的时候,它用的还是英伟达的卡,但是已经不用英伟达的生态了。V3 用英伟达的卡,但是没有用英伟达的生态,而是我们先写一个高级编译器叫 TileLang,然后基于 TileLang 的生态来完成其他所有的事情,就已经几乎不依赖英伟达的生态了。
我对国产算力是比较乐观的。我觉得在这一点上,英伟达是在掘自己的坟墓。华为的超节点,华为的 950 超节点,在性能和价格上可以完全平替英伟达的 GB200、GB300。
四张华为卡顶一张英伟达的卡。
我们跟美国在芯片上的差距,我认为生态上以后不会再有差距,但是在芯片上是四倍加两年。
我们现在主要是跟华为有合作。华为他们自己适配,但我们自己会参与这个生态,会深入参与到华为这个里面去。华为的问题还是产能不足。
我不太相信未来五年之后,我们还卡在产能的问题上。现在肯定是卡在产能问题上,今年、明年、后年,我觉得可能都还是卡在产能问题上,但五年之后,我觉得可能不一定,我还是比较乐观的。
06 竞争格局与行业判断
各家模型最终效果拉开差距,应该是一个综合上的。比较模型效果,肯定是得在相同的成本上来比较,这个才是有意义的。因为你比较两辆车,也是同价位的车来比较。
Anthropic 现在超过 OpenAI,这是不是长期的?我觉得这不是长期的,这肯定是阶段性的。OpenAI 和 Google,未来大概率还是会交替上升。
在全球 AI 中主导分工的时候,中国公司很有可能扮演的一个角色还是产量最大。常理来讲,我们的产能最大,包括芯片,芯片可能我们的产能最大,我们的电力最多。
中国人会把这产品做到最便宜,然后再在效果上,毕竟国外的商品,现在很多商品中国产跟美国产并没有太大的区别。未来可能 AI 也是这样,但是中国产的 AI 可能价格会更便宜。这个便宜可能是系统性的低,就跟其他行业中国提供的服务更便宜可能是一样的。
最终的差距应该是三方面:一方面成本,一方面时间,一方面用户体验。除此以外,可能是没有什么差距的。
成本肯定是一个差异,我觉得可能成本是排在第一位的区别。然后第二个就是时间,你什么时候能够做到。你早几个月、晚几个月,它就不一样了。
OpenAI 从一开始觉得他真的能够垄断这个世界,但是实际上他会遇到很多很多挑战者。他会遇到挑战,他就不会那么轻松。美国会遇到挑战,那么他在未来可能还会遇到中国的挑战,因为中国人愿意拿得更少,就可以给你提供这个服务。
拿得多的人会被拿得少的人打败。甚至你还不用真的拿得多,愿景如果是拿得多的话,你就会被愿景是拿得少的人给打败。其实大家都没有拿到钱,只是一个愿景。你愿景是拿得多,你就先输了,你就会面临着更大的困难。
对我们来讲,我们并不是利润要拿最多的钱,或者说算收益最大化的定价,而是只赚一个合理的收益。这是一个解释。我是相信这个事的,我并不是去为这个事情找理由,因为没必要找理由。
我觉得在很多体验方面,有可能我们是能比美国做得好的。在产品方面,产品能力上不一定会比美国差。成本应该也会比美国低,所以中国还是会有竞争力的。
成本这个很好理解,是因为他们都不用做,所以他们就不发展这个能力。他们肯定没有我们重视这个事情。我们可以把它当做一个非常重要的事情,但对于他们来讲,这个是不重要的。
大模型可能不说两家大公司、两家小公司,可能就已经比较够了。差距只有两个东西:一个是时间,一个是成本。所以不至于哪一家有暴利,我觉得不至于有暴利。成本控制得好的人就多赚一点,成本控制得差的人就少赚一点,仅此而已。
07 模型研发与技术
我们公司可能有一半的人,平时有一半的人觉得 OpenAI 是更好的。其实 Anthropic 它有先发优势,但这先发优势应该很快就没了,并不是一个它能够长期占得住的优势。大家这三家都很厉害,这三家里面的效率是最高的,它花的成本、它花掉的、它烧掉的钱应该是最少的。
多模态布局,我们一直在做。对产品来讲,它很重要;对 C 端用户产品来讲,它很重要。但是对智能的上限,它是一个组件,它不是主线本身。
我们应该会上相关的模型,就是我们 V4、V4 的后续版本会支持原生的多模态。但是我们对多模态、对智能来讲,它是个组件,我们不把它当作智能本身。
只能说语言模型的 Scaling,我现在没有看到有上限。我们现在的智力水平,或造成美国的智力水平,都没有看到上限。
我们内部很多人的想法是这样的:首先要对我们自己有用,首先是给我们自己用。然后这是实现 AGI 最快的方法。当我们自己好用,那意味着可能别人也好用,但是首先得保证我们自己好用。
我们做的模型,第一目标不是大家用得好用,而是我们自己用得好用。首先是对我们自己有用。对我们自己有用之后,我在开发下一版模型的时候就会更快。
我们叫这个叫"摸奖"。门槛很低,谁都可以去摸,但是谁能摸出什么来,这个可能我也不知道是看天赋还是看什么。所以这里并不需要我们去分配资源。只是说,我们跟其他公司不一样的地方,就是我们会花时间去讨论这个问题,会去想这个问题,然后把它当做一个重要的事情。
08 商业化与定价
我们的 API 定价是一个合理的利润,大概是我们到市场上买一批设备回来,十个月收回成本,我觉得这是一个合理的利润。
如果利润最大化,应该把价格设得更高。因为在这个价格区间,用户的需求是没有弹性的,就是我价格再翻一半,或者说我价格再抬高一倍,token 的消耗量区别不大的。
我们的一个模型,一开始我们担心需求太多,所以一开始把价格定得比较高,团队里大家不是很高兴。后来我把价格又降下来了,降到四分之一,大家就很开心。
To B 业务的上限应该还是需求,在现在这一代 AGI、AI 技术的背景下,To B 的需求应该是有限的。它会快速增长,但并不是一个无穷大的事情,最终还是受制于需求,不是算力。
我现在觉得应该是能够做到的,就都要。假如说我今年能够有几个亿美金的 B 端收入,再加上我们 C 端有用户,那么这本身就已经有一定的商业基础。以明年我们 B 端有收入,如果这个需求可以再增大的话,公司离净利润已经不远了,可能就已经是净利润了。
最坏情况卖 API,可能都能够支撑一个上市公司。就如果说技术后面没有新的进步了,我们的技术就冻结在这里了,那么最后我们就全力卖 API,把这些服务做好,我觉得也够的。
我们现在以目前的情况来看的话,我觉得最合理的做法应该是全力做通用的 Agent,其他的 Agent 优先级应该更低,包括金融、医生这些 Agent。要先做 Coding,因为 Coding Agent 能够做到很多,还有很多垂直的 Agent。现阶段我们觉得最重要的,应该还是 Coding Agent。
我觉得低成本首先是一个结果。我们的模型确实一直在模型架构上往一个更低成本的方向走,这跟我们的愿景有关系。我们还有很多在算法上的方法,成本还可以往下走。
成本往下走还有一个原因是,成本越低,我就越能训练更大的模型,我就越能承担起更大的模型。在同样算力上,在算力有限的情况下,如果我的计算效率更高,我就能够承担起更大的模型。
09 开源策略
我觉得我们是会开源的,然后我们最强的模型可能也是会开源的。因为我看不到闭源什么好处,看不到必然的好处。字节它的模型是闭源的,它有什么好处?我看不到有什么好处。
哪怕是模型开源,你把所有东西都告诉别人,这个门槛也非常高。别人要用起来,这个门槛也非常高。他要用起来,就很难;其次,他要用起来,还要成本做得很低,也很难很难,没有那么容易。
开源并不会影响收入。开源,我觉得对我们的商业模式是没有任何影响的。
我也不担心别人部署我们的模型,然后跟我们来竞争,一点都不担心。我们还希望他们能够部署起来。我们尽可能给开源社区提供帮助,协助大家能够把我们的模型部署起来。
我们在对外面打交道的时候,我们的态度是:我们只做 AGI 的主线。在对外面打交道的时候,我们是很愿意协助、帮助任何人,甚至我们的竞争对手,包括阿里、智谱、月之暗面,去做得更好。因为我们并不损失什么东西,我们本来也是开源的。
我们给的开源模型,跟我们自己部署的模型是不是一样?是一样的。我们不会说开源一个差点的模型,然后我们自己部署的时候用一个更好的模型,是不会的,是一样的。
10 数据与后训练
数据应该几乎就等于模型的一半。前面还有一个标数据的问题。我们在数据标注方面,这跟我们的资本投入有关。以我们这个资本投入的结构,支撑不起那么多高质量数据标注的成本,因为成本很高。
美国数据标注的成本跟中国数据标注成本没有什么区别。中国去标数据并没有成本优势,尤其是标高端数据上并不会有成本优势,使得我们很难投入去像美国这样标数据。这条路在中国是很难的,因为标数据实在太贵了,不管是我们外标还是我们自己标,都很难受。
现在基本上是两条腿走路。并不是说我们完全不能标,而是因为标数据有一些成本低,有一些成本高。我们先标成本低的。
你也可以认为,现在我们公司有一半的人在标数据。有一半的核心研究员,最重要的人,有一半在标数据。我们就集中在标数据。解决 AI 这个问题,在现在这个阶段靠的就是标数据。
高质量数据标注的瓶颈,我觉得是时间,就是需要时间。因为对 OpenAI 来讲、对国外来讲、对 Anthropic 来讲,他们都更早,然后资本更多,卡也更多。
大模型的幻觉问题比较影响用户的体验。幻觉问题也是有一个方法可以解决的,但是这是一个长命题。幻觉问题可以认为是一个可以通过更好的 Post-training 解决的,是一个能解、能够改善的问题。
11 组织与公司定位
首先,我们没有模仿的对象。每一步都是我们从实际情况出发,实事求是,根据实际情况来做决策,找到我们应该怎么做。所以它是一个时代的产物,或者说是现实情况的一个反映,它并不是一个模仿的结果。
我们明确是要有商业化的。我们最终还是要能活下去,我们毕竟是一个公司,政府不会给我一分钱。
我们本质上还是一个公司,只是说我们在考虑赚哪些钱、什么时候赚钱、赚多少钱、靠什么赚钱,我们有取舍。很多公司做得很伟大,因为它有一种利润以外的追求。那个追求最后不但没有影响到它的商业化,反而能让它商业化得更好。
对于合作伙伴,其实我们这个融资是精心挑选的。首先我觉得,利益是比较一致的,就是跟我们利益最一致的、对我们最没有敌意的,或者说最希望我们能够做成功的。不是所有人都希望我们能做成功的,因为我们还是损害了很多其他人的利益的。
AI 现在不缺品位和直觉,它缺的是持续学习的能力。AI 的品位和直觉没有问题。你让它写个文章,它的品位和直觉,我觉得没有什么问题。
我们希望只做一块。我觉得 AI 这个事情很大,并不需要我……我只做一块。如果聚焦,并且我认为这里的生意利益已经足够大,就如果是 AI 时代会产生很多家万亿级别的公司,我觉得我们是其中一家。
我们希望能够扶持更多的人,但是我们并没有那么多的精力。我们是有这个意愿,并且不会有利益冲突,但是我们有没有去做是另外一回事。但至少这里边是没有利益冲突的,我们是希望合作共赢的。
Liang Wenfeng / DeepSeek Investor Q&A
梁文锋投资者交流会实录
Bilingual Chinese-English Transcript (Side-by-Side)
|
中文 (Chinese) |
English |
|
码良 / Ma Liang |
码良 / Ma Liang |
|
@cxjwin / @cxjwin |
@cxjwin / @cxjwin |
|
融资 500 亿后首次开腔:梁文锋 4 小时讲透 DeepSeek(118 条实录) |
After Raising RMB 50 Billion, He Finally Speaks: Liang Wenfeng Explains DeepSeek in a 4-Hour Session (118 Verbatim Remarks) |
|
说过「不融资、不上市」的 DeepSeek,这次融了超 500 亿元,投前估值约 3675 亿元。围绕这次融资,梁文锋在一场近 4 小时的投资者交流会上回应了一切:愿景、AGI 路线图、算力、国产芯片、竞争格局、开源,共 11 个话题。以下是他的 118 条发言,按主题整理,只保留讲话内容。 |
DeepSeek, which had said it would “never raise outside funding and never go public,” has now raised more than RMB 50 billion at a pre-money valuation of roughly RMB 367.5 billion. In connection with this financing round, Liang Wenfeng addressed everything in a nearly four-hour investor Q&A session: vision, the AGI roadmap, compute, domestic chips, the competitive landscape, and open source — eleven topics in total. Below are his 118 remarks, organized by theme, retaining only the substance of what he said. |
|
01 愿景与克制 |
01 Vision and Restraint |
|
我们一开始来做这个公司,初衷没有想到说我最后要赚多少钱,要到资本市场上去,要上市,要怎么样的。最开始的几十个人完全没有这么想过,如果他这么想,他就不会来。 |
When we first started this company, it never occurred to us that we would ultimately be thinking about how much money to make, about going to the capital markets, about going public, or anything like that. The first few dozen people never thought that way at all — if they had thought that way, they wouldn’t have come. |
|
我们是怀着一个对这个世界非常大的善意来做这个事情,我们觉得这是对人类有用的,这是一个金钱以外的事情。我们出发的初衷、我们的愿景,以及我们保持到现在的这个愿景,不是按照一个商业利益最大化的方式来做的。 |
We’re doing this out of a great sense of goodwill toward the world. We feel that this is something useful to humanity, something beyond money. Our original intention, our vision, and the vision we’ve maintained to this day were never built around maximizing commercial gain. |
|
管理一个大公司,靠的不是你的规章制度,靠的是愿景。愿景不是挂在墙上的标语,愿景是你怎么做,不是怎么说,就是你怎么实际运行。 |
Running a large company depends not on your rules and regulations, but on vision. Vision isn’t a slogan hung on the wall — vision is how you actually act, not what you say, it’s how you actually operate. |
|
我们是没有组织的,就是愿景驱动的,靠一个愿景来组织。我们并不是以一个“我要实现什么 KPI、没有考核”的方式来做,只有愿景。 |
We don’t really have an “organization” in that sense — we’re vision-driven, organized around a vision. We don’t operate on the basis of “I need to hit this KPI” — there’s no performance review, only the vision. |
|
这个愿景甚至也不是成文的,并不是写出来的,没有写出来过任何东西。这个愿景是在我们做事情的方法、我们对待这个世界的态度里。 |
This vision isn’t even written down anywhere — it has never been put into writing. This vision lives in the way we do things and in our attitude toward the world. |
|
我们并没有非常多的其他优势,我们没有什么本事,我们并没有比别人有钱,也没有说我们人员比其他公司更好,其实没有的。我们两年前成立这个公司的时候,我们又没有很多钱,又没有很多卡,又没有什么知名度,又没有什么号召力,我们就是一群非常平凡的人。 |
We don’t really have many other advantages. We don’t have any special skills, we’re not richer than anyone else, and our people aren’t better than those at other companies — none of that is true. When we founded this company two years ago, we didn’t have much money, we didn’t have many GPUs, we had no name recognition, and we had no ability to rally people — we were just a group of very ordinary people. |
|
你越克制,可能就越容易做成,或者说至少到目前为止是印证的,到目前为止是能解释得通的。否则没有办法能够解释为什么我们能够做成:我们并没有什么武器,起点又非常低,资源又非常少,我们的人其实也就是随机的一群平凡的人。 |
The more restraint you exercise, the more likely you may be to succeed — or at least, that’s been borne out so far, and it’s the only thing that makes sense to me so far. Otherwise there’s no way to explain why we’ve been able to succeed: we don’t really have any “weapons,” our starting point was extremely low, our resources were extremely limited, and our people were essentially just a random group of ordinary people. |
|
AI 这个事情太大了,利益太大了。我们非常克制,只要能够做成,最后利益都会非常大。你随便分一点,利益就非常大,所以现在根本不用考虑拿这里面的哪一部分利益、怎么拿,我觉得根本不用考虑这件事情,因为这个利益足够大了。 |
AI is simply too big a thing, and the stakes are too enormous. We exercise a great deal of restraint, because as long as we succeed, the eventual payoff will be enormous. Even a small slice of it is enormous, so there’s really no need right now to think about which piece of that payoff to take, or how to take it. I don’t think we need to consider that at all, because the payoff is already big enough. |
|
我们去年春节用户突然很多,但是我们并没有去追求我要留这些用户,或者说拿这些用户来变现,或者说我要去抢这些商业利益,在用户上面兑现。我们没有去抢用户,没有去赚钱,但我们很努力想办法把用户服务好。 |
Last Chinese New Year our user numbers suddenly spiked, but we didn’t chase retaining those users, or monetizing them, or grabbing commercial value and cashing it in on our users. We didn’t fight to hold onto users or to make money from them — but we worked very hard to figure out how to serve our users well. |
|
我们并不会有这样的想法,说我要做成下一个超级 App,然后我要去跟谁竞争,我要做成下一个字节、做成下一个腾讯,完全没有这样的想法。我觉得后面的 AGI 机会应该是非常大的,后面的 AGI 机会永远是非常大的。 |
We would never think in terms of “I want to build the next super-app,” or “I need to compete with so-and-so,” or “I want to become the next ByteDance, the next Tencent” — we don’t think that way at all. I believe the AGI opportunity ahead is enormous, and the AGI opportunity will always be enormous. |
|
克制是一种战略。就在于有时候你可以舍弃一些,来换更多其他的东西。不开源这个事,其实也是一样的,也可以认为是我们的压力,也可以认为是我们的让利。 |
Restraint is itself a strategy. Sometimes you can give something up in exchange for more of something else. Not open-sourcing, for that matter, is really the same thing — you could see it as pressure on us, or you could see it as a concession we’re making. |
|
这种克制,我理解是这种克制从长远上来讲,能够增加我们做成 AGI 的概率。在考虑一件事的时候,我是毫不怀疑 AGI 会有非常大的商业价值。那么在这个基础上,我优先考虑的不是我怎么多加一点份额,我怎么多拿一些份额,我优先考虑的是我怎么增加我能够做成的概率。 |
As I understand it, this kind of restraint increases, over the long run, the probability that we’ll succeed in building AGI. When I think about any given issue, I have no doubt whatsoever that AGI will have enormous commercial value. Given that, what I prioritize isn’t how to grab a bit more market share or take a bit more of the pie — what I prioritize is how to increase the probability that we succeed. |
|
我们一直非常克制,不愿意跟任何一家互联网大厂或者小厂成为对手。我希望我能够给他赋能,或者希望我能够协助大家去做这个事情,希望能够帮助大家做这个事情。 |
We have always been very restrained, and we’re unwilling to become an adversary to any internet giant or small company. I hope I can empower them, or help everyone do this work — I hope I can help everyone accomplish this. |
|
我觉得我们之前秉持这种态度,其实我们并没有因此而少拿到任何东西,并没有因为我开源,并没有因为我们的善意或者说我对其他人提供帮助,而导致我们少拿了任何东西。反而可能还有加分。这个看起来违反直觉,但它确实是这样的。 |
I think that by holding to this attitude, we actually haven’t lost out on anything as a result. Open-sourcing, or our goodwill, or the help we’ve given others, hasn’t cost us anything. If anything, it may have earned us bonus points. This seems counterintuitive, but it’s really true. |
|
我们是以 AGI 为目标,但是我们一直在做商业化,所以我们才有 C 端的用户,才有 B 端的收入。从历史经验来看,这个策略是成功的。 |
Our goal is AGI, but we’ve been pursuing commercialization all along, which is why we have consumer users and enterprise revenue. Based on our track record, this strategy has been successful. |
|
02 AGI 路线图 |
02 AGI Roadmap |
|
如果说你能够把一个问题描述得很清楚,给它完整的上下文和指令,它已经超过人类了。但这里有个定义,有个前提是:你给它完整的上下文,给它完整的指令。 |
If you can describe a problem very clearly and give it full context and instructions, it has already surpassed humans. But there’s a definition here, a precondition: that you give it complete context and complete instructions. |
|
AI 并不能够替代你的员工。但如果说 AI 具有持续学习的能力,它跟你的员工一样,到公司学习两个月,那么就可以替代天下的人了,所以我们离下一步还差一个持续学习。 |
AI cannot replace your employees. But if AI had the ability to keep learning continuously — the way an employee would, spending two months learning at the company — then it could replace everyone in the world. So the next step we’re missing is continual learning. |
|
AI 的发展,我们可以理解成它是一个阶梯。去年走的阶梯是思维链。因为我们发现,通过思维链的方式,可以让智能达到一个更高的水平。 |
We can think of AI’s development as a series of steps. Last year’s step was chain-of-thought, because we found that chain-of-thought could bring intelligence to a higher level. |
|
今年的阶梯就是 Agent,因为我们发现,用 Agent 的方式,即便多些事情也可以做,它的能力范围会更大,它的智能上限会更高。Agent 要用到 CoT,然后 CoT 也要用到前面的阶梯,前面的阶梯就是语言模型,所以它并没有一步是白走的。 |
This year’s step is Agents, because we found that with an agent-based approach, it can handle more tasks, its range of capability is larger, and its ceiling for intelligence is higher. Agents rely on chain-of-thought, and chain-of-thought in turn relies on the step before it, which is the language model itself — so none of the steps were wasted. |
|
在 Agent 之后,我们觉得应该要解决的问题是持续学习,就是怎么让模型可以持续地学习,而不是说你要给它一个很强的训练,它应该能够像人一样做一个比较长时间的持续学习。 |
After Agents, we believe the next problem to solve is continual learning — how to let a model keep learning over time, rather than requiring intensive training all at once. It should be able to engage in long-term, continuous learning the way a person does. |
|
持续学习之后,可能我们就会来到一个奇点。这个奇点就是,当这个模型能够持续学习之后,它已经能够做人类能做的所有事情了。它就能够自己开发自己的版本,能够自己再研究,然后开发自己的下一个版本,开发更好的人工智能模型。 |
After continual learning, we may arrive at a singularity. That singularity is this: once a model can learn continuously, it will already be able to do everything a human can do. It will be able to develop its own next version by itself, to conduct its own research, and then develop its own next version — building better AI models. |
|
这个奇点,它并不是一个奇点,它也是一个渐进的过程。这个过程可能也是一个比较长的渐变,它不是个突变。但是习惯性地,我们都认为它可能是个奇点。 |
This “singularity” isn’t really a singularity — it, too, is a gradual process. That process may be a fairly long, gradual transition rather than a sudden leap. But out of habit, we all tend to call it a singularity. |
|
这是我们的推测,这是我们觉得这个时间表应该是:先解决学习的问题,然后再到那个智能的奇点,能自我迭代的奇点,然后才是具身智能。到具身智能之后,它就走进现实世界,可以给你做家务,可以给你养老。 |
This is our conjecture — this is what we think the timeline should look like: first solve the problem of learning, then arrive at the intelligence singularity — the singularity where it can iterate on itself — and only after that comes embodied intelligence. Once we reach embodied intelligence, it will step into the physical world and be able to do your housework, take care of you in your old age. |
|
如果说我们先解决持续学习,再解决那个自我迭代的奇点,再解决具身智能,这个路上就很轻松。因为到后面之后,你可以用前面的技术来帮助开发后面的技术。 |
If we first solve continual learning, then solve the self-iterating singularity, and then solve embodied intelligence, that path becomes much easier. Because by that point, you can use the earlier technology to help develop the later technology. |
|
我们只做 AGI 的主线。AI 领域很广泛,有很多东西我们觉得它不在这个主线上面,比如说 3D、视频生成,我觉得可能跟智能的主线没有太大的关系,我们不会去做。 |
We only work on the main line toward AGI. The field of AI is very broad, and there’s a lot that we don’t think belongs on that main line — 3D generation and video generation, for example. I don’t think those have much to do with the main line of intelligence, so we won’t be pursuing them. |
|
视频生成一开始出来的时候就很火,好像这是必须要做的,如果你不做,你就不是一个 AI 公司一样。所以我就很奇怪,这个其实你只要仔细去想一下,它跟智能的路线图是没有什么关系的。 |
When video generation first came out it was very hot, as if it were something every AI company had to do — as if you weren’t a real AI company if you didn’t do it. I find that strange, because if you think about it carefully, it really has nothing to do with the roadmap toward intelligence. |
|
在商业上,它是个好生意,在商业上是个好生意。但是这跟智能没有什么关系。我们不会因为它是一个好商业而去做它,我们只会因为它是智能路线图上的东西,才会去做。 |
Commercially, it’s a good business — it really is a good business. But that has nothing to do with intelligence. We won’t do something just because it’s a good business; we’ll only do it because it belongs on the roadmap toward intelligence. |
|
从我们的判断,世界模型和智能还不是现在这个阶段最重要的事情。最重要的才是 AI 训练,以及 AI 训练之后怎么解决持续学习。这是我们公司的判断,当然每个公司它的判断是不一样的。 |
In our judgment, world models aren’t yet the most important thing for intelligence at this stage. What matters most is AI training, and then, after training, how to solve continual learning. That’s our company’s judgment — of course, every company’s judgment differs. |
|
我们现在比较相信一个叙事是,AI 可以加速 AI 的研究。就是说,它不是线性的,因为你可以用 AI 来加速你自己的研究,所以它到后面可能是非线性的。 |
Right now we tend to believe in the narrative that AI can accelerate AI research itself. In other words, progress isn’t linear — because you can use AI to accelerate your own research, which means it may become nonlinear down the road. |
|
我觉得具身肯定还是要进入,最终具身。因为对一个正常人来讲,他的需求并不是电脑,对不对?因为正常的人,他吃喝玩乐、衣食住行,他不需要电脑。他需要的是,所以他还是需要具身智能来解决具体人类的需求。 |
I think embodiment definitely still needs to happen — ultimately we need embodiment. Because for an ordinary person, what they need isn’t a computer, right? An ordinary person’s needs — eating, drinking, entertainment, clothing, food, housing, transportation — don’t require a computer. What they need is embodied intelligence to address concrete human needs. |
|
我们希望 AGI 能做什么呢?它能帮我迭代下一版模型,能帮我迭代下一版模型一样。如果有了具身之后,我们希望它做的也是,让它来迭代下一版的具身,它来做下一版机器人。 |
What do we hope AGI will be able to do? We hope it can help us iterate on the next version of the model — help us iterate on the next model version. And once we have embodiment, what we’d want it to do is the same: let it iterate on the next generation of embodiment, let it build the next version of the robot. |
|
下一代模型的核心能力,必须要有持续学习的能力,它才能叫下一代模型。在那之前,我们能做的就是降低成本,然后效果做得更好,速度做得更快。但是要有大突破,它应该是具备持续学习的。 |
The core capability of the next-generation model must be continual learning — only then can it truly be called a next-generation model. Before that, what we can do is lower costs, improve performance, and increase speed. But for a real breakthrough, it needs to have the ability to learn continuously. |
|
现在的 Agent 的能力受限,是因为它不能持续学习,它不能有效地持续学习。如果说能够先把持续学习做完,那 AI 的能力是非常强的,它能够非常大地提升我们自己研究的效率。 |
Today’s agents are limited because they can’t learn continuously — they can’t effectively engage in continual learning. If we can solve continual learning first, AI’s capability will become extremely strong, and it will greatly improve the efficiency of our own research. |
|
持续学习先做出来,通用智能可能就很容易了,用它来做就很容易。所以我说这是一个我们比较希望看到的结果,我们比较省力,我们就轻松。否则现在你要去人工做通用智能,它是一个比较累、比较苦,是一个数据密集、人力密集的事情,性价比也不高。 |
If we solve continual learning first, general intelligence may become quite easy to achieve using it. That’s why I say this is an outcome we’d really like to see — it would save us a lot of effort and make things easier. Otherwise, building general intelligence manually right now is exhausting and grueling work — it’s data-intensive, labor-intensive, and not very cost-effective. |
|
03 团队与人才 |
03 Team and Talent |
|
我们前面的经历给我的启示是,AGI 这个愿景是很强大的。这个人才优势不是说我的人比他更聪明,而是这些人才我怎么组织起来,怎么激励他,然后怎么合作。 |
What our earlier experience has taught me is that the AGI vision is extremely powerful. Our talent advantage isn’t that our people are smarter than anyone else’s — it’s about how we organize this talent, how we motivate them, and how we get them to collaborate. |
|
你把聪明的人聚在一起,并不是说他自然而然就能够合作,自然而然就能够非常有激情地去奔一个目标、去完成的,所以你需要一个愿景。 |
Putting smart people together doesn’t mean they will naturally cooperate, or naturally be passionately driven toward a single goal and get it done. That’s why you need a vision. |
|
我们最大的核心利益是要保持团队的稳定性。这是我们最大的核心利益,甚至可以认为是唯一的核心利益。只要我能够保持团队的稳定性,我一定能做成,一定能做成 AGI,就这么简单。 |
Our single greatest core interest is maintaining the stability of the team. This is our greatest core interest — you could even say it’s our only core interest. As long as I can keep the team stable, I will definitely succeed — I will definitely achieve AGI. It’s that simple. |
|
钱肯定不是问题,资源不是问题,其他要素都是容易获得的。对我们来讲,只有一个核心利益,只有一个没法退让的:我们必须要保持团队的稳定性。 |
Money definitely isn’t the issue, resources aren’t the issue — all the other factors are easy to obtain. For us, there’s only one core interest, only one thing we cannot compromise on: we must maintain the stability of the team. |
|
这也是我们面临的一个非常大的挑战,或者说,我觉得是最大的风险。当然,这个风险随着我们近期的这一次融资,得到了比较大的解除。因为大家拿到的期权都还是比较多的,金额还是比较大的。 |
This is also a very major challenge we face — or, I’d say, it’s our greatest risk. Of course, this risk has been substantially alleviated by our recent financing round, because everyone has received a fairly large amount of stock options, worth a fairly substantial sum. |
|
从团队稳定性来讲,只要最重要的一些员工、最老的员工能够稳定,那么其他人是不太会走的。其他人哪怕期权少一点、收入少一点,他也不会走。因为他不是全奔着钱来的,大家都希望在一个能够做成 AGI 的环境里面去做这个事情。 |
In terms of team stability, as long as the most important employees, the longest-tenured employees, remain stable, the others are unlikely to leave. Others won’t leave even with fewer options or lower income, because they didn’t come here purely for the money — everyone wants to work in an environment where AGI can actually be achieved. |
|
其他都是时间问题,其他最多导致我们晚半年、晚一年,但是不会说做不出来。肯定是不缺钱,肯定是不缺资源,其实这些都是不缺的。 |
Everything else is just a matter of time. At most, other factors might delay us by six months or a year, but they won’t stop us from getting there. We definitely aren’t short of money, and we definitely aren’t short of resources — none of that is actually in short supply. |
|
我们跟美国的差距主要在资源上面,然后人上面差距不是很大的。人上面几乎没有差距,因为就是同一批人,可能是中国人。中国人出去的时候,有一些人留在国内,有些人留在国外,有些人去国外,他并没有说是聪明的人去国外,没有的。 |
Our gap with the U.S. is mainly in resources — the gap in people isn’t very large. There’s almost no gap in talent, because it’s essentially the same pool of people, many of them Chinese. Among Chinese people who go abroad, some stay in China and some go overseas — it isn’t the case that the smarter ones go overseas. That’s not true. |
|
人才不是瓶颈,资源是最大的瓶颈。资源首先影响到人才培养,因为算力少,我们能做的实验机会比较少,所以我们的人才整体上比美国有差距。人才的差距,本质上也是因为算力的差距。 |
Talent isn’t the bottleneck — resources are the biggest bottleneck. Resources affect talent development first, because with less compute, we have fewer opportunities to run experiments, which means our talent overall lags behind that of the U.S. The gap in talent is, fundamentally, also a gap in compute. |
|
AI 人才的短缺也是阶段性的,并且我们已经看到,大幅度被缓解了。因为 AI 人真的不缺,每个公司很快会把人培养出来,培养人是很快的。 |
The shortage of AI talent is also just a phase, and we’ve already seen it be substantially alleviated. AI talent really isn’t in short supply — every company will quickly train up its own people; developing talent happens fast. |
|
国内现在做模型的公司有点太多了,还是太多了。美国可能就三家,中国做基模的东西太多了。最终一定是不需要那么多人去做基模的,一定会收敛。 |
There are a bit too many companies in China building models right now — really too many. The U.S. might have just three, while China has far too many companies building foundation models. Ultimately, there’s no need for that many players building foundation models — the field will inevitably consolidate. |
|
我们公司的管理其实是两条线:一条线是从上到下,一条是从下到上。从下而上,就是每个人自己想做什么,自己做,没有人管他,没有 KPI。 |
Management at our company actually runs along two lines: one top-down, one bottom-up. Bottom-up means everyone decides what they want to work on and does it themselves — no one manages them, there are no KPIs. |
|
一般我们希望员工还有一半的时间,是不被安排的,他想做什么就做什么。这是一个研究的范围,让他可以自己去探索,按照他觉得什么重要,他去探索什么,没有前置的要求。 |
Generally, we want employees to have about half their time unscheduled, free to work on whatever they want. This is meant as room for research, letting them explore on their own — they explore whatever they think is important, with no prior requirements. |
|
我们一般也不太加班。加班有两个原因。第一个是,做研究是需要一个比较松弛的环境。你如果逼得很紧,就没法做研究。因为既然就是要你自己有这个兴趣,你自己平时要去想这些问题,所以得是在一个比较松弛的环境里,才有可能能够探索。 |
We also generally don’t work much overtime. There are two reasons for this. First, research requires a relatively relaxed environment. If you push people too hard, they can’t do research, because research depends on your own genuine interest — you need to be thinking about these problems on your own time, which requires a relatively relaxed environment in order to be able to explore. |
|
第二个是,我们非常聚焦。我们非常聚焦,就意味着我们要做的事情很少。那我就没那么多事情要做,我就不需要加班。这个跟前面的克制是一脉相承的。 |
The second reason is that we’re extremely focused. Being extremely focused means there are very few things we actually need to do. So I don’t have that much to do, and I don’t need to work overtime. This is consistent with the restraint I mentioned earlier. |
|
我们公司整体上是建立在共识的基础上的,我并不是说我一个人决定所有事情,而是我要寻求共识。我在公司内的权威以及在公司内的影响力,是建立在共识的基础上的。 |
Our company as a whole is built on the foundation of consensus. It’s not that I alone decide everything — rather, I seek consensus. My authority within the company, and my influence within the company, are built on the foundation of consensus. |
|
这个决策机制其实是一种寻求共识的机制,这并不是说我能够推动一个什么事情,它一定是共识,我才能够推得下去,然后我才会去推。 |
This decision-making mechanism is really a consensus-seeking mechanism. It’s not that I can simply push something through on my own — it has to be a consensus before I can push it forward, and only then will I push it. |
|
随着人员的增加,我们会做这个调整。应该是我们马上就得做这个调整,因为我已经在做这个调整。已经不做这个调整的话,很多事情是没法推进下去的。确实有很多部门,是应该有组织架构的。 |
As headcount grows, we will make adjustments to this. In fact, we need to make this adjustment right away, because I’m already making it. Without this adjustment, a lot of things simply can’t move forward. There really are many departments that should have a proper organizational structure. |
|
04 算力与资源 |
04 Compute and Resources |
|
我们需要多少卡?现在肯定是越多越好。在我们能够承受的范围内,肯定是卡越多越好,这是毫无疑问的。所以我们现在策略是,在合理的价格里面,能买到多少卡就买多少卡。 |
How many GPUs do we need? Right now, more is definitely better. Within what we can afford, more GPUs is unquestionably better — there’s no doubt about that. So our current strategy is to buy as many GPUs as we can at a reasonable price. |
|
实际上,要把这么多钱花完非常不容易,买不到那么多卡,很难买,而且价格也很高,也不能说花非常高的价格去买,还得确保这个价格是合理的。如果今年能够花掉两百亿,那么就属于我们的采购部门业绩超级好了。 |
In reality, spending that much money isn’t easy at all — you can’t buy that many GPUs, they’re hard to get, and prices are high. We also can’t just pay an extremely high price for them; we still need to make sure the price is reasonable. If we can spend RMB 20 billion this year, that would already mean our procurement team has performed exceptionally well. |
|
我们跟美国之间最大的差距是在资源上面。算力资源一方面是国内本身卡就买不到,另外一方面是我们的资本投入比美国要少。我们在资本投入层面上少了很多,基于人才的工资在这里面占的比重是很低的。你看他们开的薪水一个亿美金这种,但算起来,人才的薪水还是占比很少,大头还是算力。 |
Our biggest gap with the U.S. is in resources. On the compute side, one issue is that we simply can’t buy GPUs domestically, and another is that our capital investment is smaller than that of the U.S. We’re investing far less capital, and the proportion accounted for by talent salaries within that is actually quite small. You see salaries like $100 million being offered over there, but when you actually calculate it, talent compensation still accounts for a small share — the bulk of it is compute. |
|
我们看到的所有区别,包括人才的区别、模型能力的区别、应用的区别,可以认为都是因为算力资源上的区别。 |
All the differences we see — differences in talent, differences in model capability, differences in applications — can all be attributed to differences in compute resources. |
|
我们跟美国的差距可能是落后美国 12 个月,落后美国可能 12 到 18 个月,或者说 6 到 12 个月。反正简单说,就是落后美国两年,然后只用美国二十分之一的算力把这个事情做出来。 |
Our gap with the U.S. might be about 12 months behind, or maybe 12 to 18 months, or 6 to 12 months. Simply put, we’re roughly two years behind the U.S., and we’ve achieved this using only one-twentieth of the compute the U.S. uses. |
|
这个叙事就是落后一到两年,但是只用它二十分之一的算力。那么未来我们要把这个叙事改写,就是我们用它几分之一的算力,但是把这个时间缩得更短,缩到 6 个月、3 个月,我觉得这是一个目标。 |
That’s the narrative right now — one to two years behind, but using only a twentieth of the compute. Going forward, we want to rewrite that narrative: using some fraction of the compute, but shrinking that time gap down to six months, or three months. I think that’s a goal worth pursuing. |
|
Scaling,我们是信 Scaling,肯定是规模越大,效果越好,能够解锁更多的功能。阻止我们 Scaling 的其实就是算力,并不是我们不想 Scaling,是我们没有那么多的算力去做这个 Scaling。 |
On scaling — we believe in scaling. Bigger scale definitely means better results and unlocks more capabilities. What’s actually holding us back from scaling is compute. It’s not that we don’t want to scale; it’s that we don’t have enough compute to do it. |
|
我们训练这么大的模型,并不是因为我觉得这么大的模型就够了,而是我刚好有这么多资源。我是按照我的资源来算,我能够接受、能够训练的模型是多大,是这样算出来的,并不是这个模型就够了。 |
We train models this large not because we think this size is enough, but because that’s simply how many resources we happen to have. I calculate the size of model I can afford to accept and train based on my resources — that’s how I arrive at it, not because this model size is sufficient. |
|
硅谷在说 Scaling 到头的时候,那是对硅谷来说;对中国人来讲,我们离那个还很远,我们根本就没有 Scaling 到那个程度。这个 Scaling 包括数据的 Scaling、模型规模的 Scaling,然后训练成本。 |
When Silicon Valley talks about scaling having hit its limits, that’s true for Silicon Valley. For us in China, we’re still very far from that — we haven’t scaled anywhere near that level at all. This scaling includes data scaling, model-size scaling, and training cost. |
|
05 国产芯片与生态 |
05 Domestic Chips and Ecosystem |
|
英伟达 CUDA 的护城河在快速地被瓦解。一方面是现在有了 AI,然后有了 AI 之后,我要建立起这个生态比以前容易很多了,因为 AI 可以写代码。 |
Nvidia’s CUDA moat is being eroded rapidly. On one hand, now that we have AI, building up an ecosystem has become much easier than before, because AI can write code. |
|
现在计算卡的市场已经比游戏卡更大了,就没有理由这两个还需要耦合起来。现在的趋势是,以后就不再耦合了。那么专用芯片,不管是华为还是英伟达自己,以后都是专用芯片,都不是之前的这些东西了。 |
The compute-card market is now bigger than the gaming-card market, so there’s no longer any reason for the two to remain coupled together. The current trend is that they will no longer be coupled going forward. So dedicated chips — whether from Huawei or from Nvidia itself — will all become purpose-built chips going forward, no longer the kind of thing they were before. |
|
国产 AI 芯片替代现在是有个历史性机会的。我们认为,未来一年之内,我们能够看到有一个事情被验证:国产芯片的生态完全没有问题。之前认为是有问题的,认为是用不起来、不好用,但是未来我觉得一年之内,我们能够扭转这个认知,或者会用事实来扭转这些。 |
There is now a historic opportunity for domestic AI chips to achieve substitution. We believe that within the next year, we will be able to see one thing validated: that there is nothing wrong at all with the domestic chip ecosystem. It used to be thought that there were problems — that the chips couldn’t really be used, that they weren’t good to use — but I think that within a year, we’ll be able to reverse that perception, or the facts will reverse it for us. |
|
国产 AI 芯片的硬件和生态都没有问题,唯一有问题的是产能不够。国产卡适配这一点,没有障碍,英伟达挡不住。如果是在一个正常的商业环境里面,我能买到英伟达的卡,那么国产替代是比较难的;但是在英伟达的卡买不到的情况下,所有人都迫不得已,都要去做国产芯片。 |
There’s nothing wrong with the hardware or the ecosystem of domestic AI chips — the only problem is insufficient production capacity. As for adapting to domestic chips, there’s no obstacle there; Nvidia can’t stop that. In a normal commercial environment, where I could just buy Nvidia chips, domestic substitution would be relatively difficult. But when Nvidia chips can’t be bought, everyone is forced to turn to domestic chips. |
|
V3 训练的时候,它用的还是英伟达的卡,但是已经不用英伟达的生态了。V3 用英伟达的卡,但是没有用英伟达的生态,而是我们先写一个高级编译器叫 TileLang,然后基于 TileLang 的生态来完成其他所有的事情,就已经几乎不依赖英伟达的生态了。 |
When training V3, we still used Nvidia GPUs, but we had already stopped using Nvidia’s ecosystem. V3 used Nvidia’s hardware, but not Nvidia’s software ecosystem — instead we first wrote a high-level compiler called TileLang, and then built everything else on top of the TileLang ecosystem, so we had already become almost independent of Nvidia’s ecosystem. |
|
我对国产算力是比较乐观的。我觉得在这一点上,英伟达是在掘自己的坟墓。华为的超节点,华为的 950 超节点,在性能和价格上可以完全平替英伟达的 GB200、GB300。 |
I’m fairly optimistic about domestic compute. I think on this front, Nvidia is digging its own grave. Huawei’s super nodes — Huawei’s 950 super node — can fully substitute for Nvidia’s GB200 and GB300 in terms of both performance and price. |
|
四张华为卡顶一张英伟达的卡。 |
Four Huawei chips are equivalent to one Nvidia chip. |
|
我们跟美国在芯片上的差距,我认为生态上以后不会再有差距,但是在芯片上是四倍加两年。 |
As for our gap with the U.S. in chips, I believe there will no longer be a gap in the ecosystem going forward, but on the chip itself, the gap is a factor of four plus two years. |
|
我们现在主要是跟华为有合作。华为他们自己适配,但我们自己会参与这个生态,会深入参与到华为这个里面去。华为的问题还是产能不足。 |
Right now we mainly work with Huawei. Huawei does its own adaptation work, but we also participate in this ecosystem ourselves — we get deeply involved with Huawei’s efforts. Huawei’s problem is still insufficient production capacity. |
|
我不太相信未来五年之后,我们还卡在产能的问题上。现在肯定是卡在产能问题上,今年、明年、后年,我觉得可能都还是卡在产能问题上,但五年之后,我觉得可能不一定,我还是比较乐观的。 |
I don’t really believe that five years from now we’ll still be stuck on the production-capacity problem. Right now we’re definitely stuck on capacity — this year, next year, the year after, I think we’ll probably still be stuck on capacity — but five years from now, I think that may no longer be the case. I remain fairly optimistic. |
|
06 竞争格局与行业判断 |
06 Competitive Landscape and Industry Outlook |
|
各家模型最终效果拉开差距,应该是一个综合上的。比较模型效果,肯定是得在相同的成本上来比较,这个才是有意义的。因为你比较两辆车,也是同价位的车来比较。 |
The ultimate gap in performance among different models should be judged comprehensively. Comparing model performance only makes sense if you compare them at the same cost — after all, when you compare two cars, you compare cars in the same price range. |
|
Anthropic 现在超过 OpenAI,这是不是长期的?我觉得这不是长期的,这肯定是阶段性的。OpenAI 和 Google,未来大概率还是会交替上升。 |
Is Anthropic’s current lead over OpenAI a long-term one? I don’t think it is — I’m sure it’s just a phase. OpenAI and Google will most likely keep taking turns rising in the future. |
|
在全球 AI 中主导分工的时候,中国公司很有可能扮演的一个角色还是产量最大。常理来讲,我们的产能最大,包括芯片,芯片可能我们的产能最大,我们的电力最多。 |
When it comes to the global division of labor in AI, the role Chinese companies are most likely to play is that of the largest producer. Logically speaking, our production capacity is the largest — including chips; our chip production capacity may be the largest, and we have the most electricity. |
|
中国人会把这产品做到最便宜,然后再在效果上,毕竟国外的商品,现在很多商品中国产跟美国产并没有太大的区别。未来可能 AI 也是这样,但是中国产的 AI 可能价格会更便宜。这个便宜可能是系统性的低,就跟其他行业中国提供的服务更便宜可能是一样的。 |
Chinese companies will make this product the cheapest, and then compete on performance as well — after all, for many foreign goods today, there isn’t much difference between products made in China and made in the U.S. In the future, AI may be the same way, except AI made in China may be priced lower. That lower price may be a systemic one, similar to how China provides cheaper services across other industries. |
|
最终的差距应该是三方面:一方面成本,一方面时间,一方面用户体验。除此以外,可能是没有什么差距的。 |
The ultimate gap should come down to three things: cost, timing, and user experience. Beyond that, there may not be much of a gap at all. |
|
成本肯定是一个差异,我觉得可能成本是排在第一位的区别。然后第二个就是时间,你什么时候能够做到。你早几个月、晚几个月,它就不一样了。 |
Cost is definitely one difference — I think cost is probably the number one difference. The second is timing: when you’re able to get something done. Being a few months earlier or a few months later makes a real difference. |
|
OpenAI 从一开始觉得他真的能够垄断这个世界,但是实际上他会遇到很多很多挑战者。他会遇到挑战,他就不会那么轻松。美国会遇到挑战,那么他在未来可能还会遇到中国的挑战,因为中国人愿意拿得更少,就可以给你提供这个服务。 |
OpenAI believed from the start that it could truly monopolize the world, but in reality it will run into a great many challengers. Once it faces challenges, things won’t be so easy for it. The U.S. will face challenges, and going forward it may also face challenges from China, because Chinese companies are willing to take less and still provide the service. |
|
拿得多的人会被拿得少的人打败。甚至你还不用真的拿得多,愿景如果是拿得多的话,你就会被愿景是拿得少的人给打败。其实大家都没有拿到钱,只是一个愿景。你愿景是拿得多,你就先输了,你就会面临着更大的困难。 |
Whoever takes more will be beaten by whoever takes less. In fact, you don’t even need to actually take more — if your vision is to take more, you’ll be beaten by someone whose vision is to take less. In truth, no one has actually gotten the money yet — it’s just a vision. If your vision is to take more, you’ve already lost, and you’ll face greater difficulties. |
|
对我们来讲,我们并不是利润要拿最多的钱,或者说算收益最大化的定价,而是只赚一个合理的收益。这是一个解释。我是相信这个事的,我并不是去为这个事情找理由,因为没必要找理由。 |
For us, our approach isn’t to maximize profit or price for maximum returns — it’s to earn only a reasonable return. That’s an explanation, and I genuinely believe in it; I’m not making up a justification for it, because there’s no need to. |
|
我觉得在很多体验方面,有可能我们是能比美国做得好的。在产品方面,产品能力上不一定会比美国差。成本应该也会比美国低,所以中国还是会有竞争力的。 |
I think in many aspects of user experience, we may well be able to do better than the U.S. On the product side, our product capability won’t necessarily be worse than in the U.S. Costs should also be lower than in the U.S., so China will still be competitive. |
|
成本这个很好理解,是因为他们都不用做,所以他们就不发展这个能力。他们肯定没有我们重视这个事情。我们可以把它当做一个非常重要的事情,但对于他们来讲,这个是不重要的。 |
The cost issue is easy to understand — it’s because they don’t need to worry about it, so they never develop that capability. They definitely don’t care about this the way we do. We treat it as something very important, but for them, it just isn’t important. |
|
大模型可能不说两家大公司、两家小公司,可能就已经比较够了。差距只有两个东西:一个是时间,一个是成本。所以不至于哪一家有暴利,我觉得不至于有暴利。成本控制得好的人就多赚一点,成本控制得差的人就少赚一点,仅此而已。 |
For large models, having a couple of big companies and a couple of small companies might already be enough. The gap comes down to only two things: time and cost. So no single company is going to earn windfall profits — I don’t think it will get to that point. Whoever controls costs well will earn a bit more, and whoever controls costs poorly will earn a bit less. That’s really all there is to it. |
|
07 模型研发与技术 |
07 Model Development and Technology |
|
我们公司可能有一半的人,平时有一半的人觉得 OpenAI 是更好的。其实 Anthropic 它有先发优势,但这先发优势应该很快就没了,并不是一个它能够长期占得住的优势。大家这三家都很厉害,这三家里面的效率是最高的,它花的成本、它花掉的、它烧掉的钱应该是最少的。 |
Maybe half the people at our company generally think OpenAI is better. Anthropic really does have a first-mover advantage, but that advantage should disappear soon — it’s not something it can hold onto for the long term. All three of these companies are very strong, and among the three, efficiency is highest, with the least money spent, or burned. |
|
多模态布局,我们一直在做。对产品来讲,它很重要;对 C 端用户产品来讲,它很重要。但是对智能的上限,它是一个组件,它不是主线本身。 |
We’ve been working on multimodal capabilities all along. It matters a great deal for the product — it matters a great deal for consumer-facing products. But when it comes to the ceiling of intelligence, it’s a component, not the main line itself. |
|
我们应该会上相关的模型,就是我们 V4、V4 的后续版本会支持原生的多模态。但是我们对多模态、对智能来讲,它是个组件,我们不把它当作智能本身。 |
We will likely release related models — V4 and versions after V4 will support native multimodality. But as far as multimodality and intelligence are concerned, it’s a component; we don’t treat it as intelligence itself. |
|
只能说语言模型的 Scaling,我现在没有看到有上限。我们现在的智力水平,或造成美国的智力水平,都没有看到上限。 |
I can only say that, when it comes to scaling language models, I don’t currently see a ceiling. Neither our current level of intelligence, nor the level in the U.S., shows any sign of a ceiling. |
|
我们内部很多人的想法是这样的:首先要对我们自己有用,首先是给我们自己用。然后这是实现 AGI 最快的方法。当我们自己好用,那意味着可能别人也好用,但是首先得保证我们自己好用。 |
Many people internally think this way: it first has to be useful to ourselves — first for our own use. And that’s the fastest way to achieve AGI. If it’s good for us to use, that likely means it will be good for others to use too, but the priority is making sure it works well for us first. |
|
我们做的模型,第一目标不是大家用得好用,而是我们自己用得好用。首先是对我们自己有用。对我们自己有用之后,我在开发下一版模型的时候就会更快。 |
For the models we build, the first goal isn’t for everyone else to find them good to use — it’s for us ourselves to find them good to use. It has to be useful to us first. Once it’s useful to us, I’ll be able to develop the next version of the model even faster. |
|
我们叫这个叫“摸奖”。门槛很低,谁都可以去摸,但是谁能摸出什么来,这个可能我也不知道是看天赋还是看什么。所以这里并不需要我们去分配资源。只是说,我们跟其他公司不一样的地方,就是我们会花时间去讨论这个问题,会去想这个问题,然后把它当做一个重要的事情。 |
We call this “drawing lots.” The barrier to entry is very low — anyone can take a shot — but who ends up drawing something good out of it, I honestly don’t know whether that comes down to talent or something else. So this isn’t really something we need to allocate resources for specifically. It’s just that where we differ from other companies is that we spend time discussing this issue, thinking about it, and treating it as something important. |
|
08 商业化与定价 |
08 Commercialization and Pricing |
|
我们的 API 定价是一个合理的利润,大概是我们到市场上买一批设备回来,十个月收回成本,我觉得这是一个合理的利润。 |
Our API pricing is set to earn a reasonable margin — roughly, if we go out and buy a batch of equipment, we’d recover the cost in about ten months. I think that’s a reasonable profit margin. |
|
如果利润最大化,应该把价格设得更高。因为在这个价格区间,用户的需求是没有弹性的,就是我价格再翻一半,或者说我价格再抬高一倍,token 的消耗量区别不大的。 |
If we wanted to maximize profit, we should set prices higher. Because within this price range, user demand is inelastic — even if I raised the price by half again, or doubled it, token consumption wouldn’t change much. |
|
我们的一个模型,一开始我们担心需求太多,所以一开始把价格定得比较高,团队里大家不是很高兴。后来我把价格又降下来了,降到四分之一,大家就很开心。 |
For one of our models, we were initially worried demand would be too high, so we set the price relatively high at first, and the team wasn’t very happy about it. Later I brought the price back down, cutting it to a quarter of what it was, and everyone was happy. |
|
To B 业务的上限应该还是需求,在现在这一代 AGI、AI 技术的背景下,To B 的需求应该是有限的。它会快速增长,但并不是一个无穷大的事情,最终还是受制于需求,不是算力。 |
The ceiling for our enterprise (To B) business should still be demand. Given the current generation of AGI and AI technology, To B demand should be limited. It will grow rapidly, but it isn’t an infinite thing — ultimately it’s constrained by demand, not by compute. |
|
我现在觉得应该是能够做到的,就都要。假如说我今年能够有几个亿美金的 B 端收入,再加上我们 C 端有用户,那么这本身就已经有一定的商业基础。以明年我们 B 端有收入,如果这个需求可以再增大的话,公司离净利润已经不远了,可能就已经是净利润了。 |
I now think we should be able to achieve both — we want it all. Suppose this year we can generate a few hundred million dollars in enterprise revenue, and we also have our consumer users — that alone would already give us a solid commercial foundation. If our enterprise revenue continues next year and demand keeps growing, the company won’t be far from net profit — it might already be at net profit. |
|
最坏情况卖 API,可能都能够支撑一个上市公司。就如果说技术后面没有新的进步了,我们的技术就冻结在这里了,那么最后我们就全力卖 API,把这些服务做好,我觉得也够的。 |
In the worst case, just selling API access could still support a publicly listed company. If our technology stopped advancing from here and got frozen at its current level, we could simply focus entirely on selling API access and doing that service well — I think that would still be enough. |
|
我们现在以目前的情况来看的话,我觉得最合理的做法应该是全力做通用的 Agent,其他的 Agent 优先级应该更低,包括金融、医生这些 Agent。要先做 Coding,因为 Coding Agent 能够做到很多,还有很多垂直的 Agent。现阶段我们觉得最重要的,应该还是 Coding Agent。 |
Given how things stand right now, I think the most sensible approach is to focus all our effort on general-purpose agents, with other agents given lower priority — including agents for finance or medicine. We should prioritize coding first, because coding agents can accomplish a great deal, and there are many vertical agents beyond that. At this stage, we believe the most important thing is still the coding agent. |
|
我觉得低成本首先是一个结果。我们的模型确实一直在模型架构上往一个更低成本的方向走,这跟我们的愿景有关系。我们还有很多在算法上的方法,成本还可以往下走。 |
I think low cost is, first and foremost, a result. Our models have indeed consistently moved their architecture in the direction of lower cost, and that’s connected to our vision. We still have many algorithmic approaches available, so costs can go down even further. |
|
成本往下走还有一个原因是,成本越低,我就越能训练更大的模型,我就越能承担起更大的模型。在同样算力上,在算力有限的情况下,如果我的计算效率更高,我就能够承担起更大的模型。 |
There’s another reason to push costs down: the lower the cost, the larger a model I can afford to train, the larger a model I can take on. With the same amount of compute — under limited compute — if my computational efficiency is higher, I can afford to train a larger model. |
|
09 开源策略 |
09 Open-Source Strategy |
|
我觉得我们是会开源的,然后我们最强的模型可能也是会开源的。因为我看不到闭源什么好处,看不到必然的好处。字节它的模型是闭源的,它有什么好处?我看不到有什么好处。 |
I think we will keep open-sourcing, and even our strongest model will probably be open-sourced too, because I don’t see any benefit to closed-sourcing — I don’t see any necessary benefit. ByteDance’s model is closed-source — what benefit does that get them? I don’t see any benefit. |
|
哪怕是模型开源,你把所有东西都告诉别人,这个门槛也非常高。别人要用起来,这个门槛也非常高。他要用起来,就很难;其次,他要用起来,还要成本做得很低,也很难很难,没有那么容易。 |
Even with an open-source model, even if you tell others everything, the barrier to entry remains extremely high. For others to actually put it to use, that barrier remains extremely high. It’s very hard for them to actually deploy and use it; and beyond that, it’s very, very hard for them to get the cost down low enough — it’s not that easy. |
|
开源并不会影响收入。开源,我觉得对我们的商业模式是没有任何影响的。 |
Open-sourcing doesn’t hurt revenue. I don’t think open-sourcing has any impact on our business model at all. |
|
我也不担心别人部署我们的模型,然后跟我们来竞争,一点都不担心。我们还希望他们能够部署起来。我们尽可能给开源社区提供帮助,协助大家能够把我们的模型部署起来。 |
I’m also not worried about others deploying our model and then competing with us — not worried at all. We actually hope they can get it deployed. We do our best to help the open-source community, assisting everyone in getting our models up and running. |
|
我们在对外面打交道的时候,我们的态度是:我们只做 AGI 的主线。在对外面打交道的时候,我们是很愿意协助、帮助任何人,甚至我们的竞争对手,包括阿里、智谱、月之暗面,去做得更好。因为我们并不损失什么东西,我们本来也是开源的。 |
When it comes to dealing with the outside world, our attitude is this: we only work on the main line toward AGI. In our dealings with the outside world, we’re very willing to assist and help anyone — even our competitors, including Alibaba, Zhipu, and Moonshot AI — to do better. Because we don’t actually lose anything by doing so; we were already open-source anyway. |
|
我们给的开源模型,跟我们自己部署的模型是不是一样?是一样的。我们不会说开源一个差点的模型,然后我们自己部署的时候用一个更好的模型,是不会的,是一样的。 |
Is the open-source model we release the same as the model we deploy ourselves? Yes, it’s the same. We wouldn’t open-source a lesser model and then use a better one for our own deployment — we would never do that. They’re the same. |
|
10 数据与后训练 |
10 Data and Post-Training |
|
数据应该几乎就等于模型的一半。前面还有一个标数据的问题。我们在数据标注方面,这跟我们的资本投入有关。以我们这个资本投入的结构,支撑不起那么多高质量数据标注的成本,因为成本很高。 |
Data probably accounts for almost half of what makes a model. Before that, there’s also the issue of data annotation. Our approach to data annotation is tied to our capital investment. Given the structure of our capital investment, we simply can’t support the cost of that much high-quality data annotation, because the cost is very high. |
|
美国数据标注的成本跟中国数据标注成本没有什么区别。中国去标数据并没有成本优势,尤其是标高端数据上并不会有成本优势,使得我们很难投入去像美国这样标数据。这条路在中国是很难的,因为标数据实在太贵了,不管是我们外标还是我们自己标,都很难受。 |
There’s not much difference between the cost of data annotation in the U.S. and in China. Annotating data in China doesn’t give us a cost advantage — especially not for annotating high-end data — which makes it very difficult for us to invest in data annotation the way the U.S. does. This path is a hard one in China, because annotating data is simply too expensive — whether we outsource it or do it ourselves, it’s painful either way. |
|
现在基本上是两条腿走路。并不是说我们完全不能标,而是因为标数据有一些成本低,有一些成本高。我们先标成本低的。 |
Right now, we’re basically walking on two legs. It’s not that we can’t do annotation at all — it’s that some data is cheap to annotate and some is expensive. We annotate the cheap ones first. |
|
你也可以认为,现在我们公司有一半的人在标数据。有一半的核心研究员,最重要的人,有一半在标数据。我们就集中在标数据。解决 AI 这个问题,在现在这个阶段靠的就是标数据。 |
You could also say that right now, half the people at our company are doing data annotation. Half of our core researchers — our most important people — are annotating data. We are concentrating on data annotation. At this stage, solving the AI problem really does depend on data annotation. |
|
高质量数据标注的瓶颈,我觉得是时间,就是需要时间。因为对 OpenAI 来讲、对国外来讲、对 Anthropic 来讲,他们都更早,然后资本更多,卡也更多。 |
I think the bottleneck for high-quality data annotation is time — it simply requires time. Because compared to OpenAI, compared to companies abroad, compared to Anthropic, they all started earlier, and they have more capital and more GPUs. |
|
大模型的幻觉问题比较影响用户的体验。幻觉问题也是有一个方法可以解决的,但是这是一个长命题。幻觉问题可以认为是一个可以通过更好的 Post-training 解决的,是一个能解、能够改善的问题。 |
The hallucination problem in large models significantly affects user experience. There is a way to address the hallucination problem, but it’s a long-term challenge. Hallucination can be seen as a problem that can be addressed and improved through better post-training. |
|
11 组织与公司定位 |
11 Organization and Company Positioning |
|
首先,我们没有模仿的对象。每一步都是我们从实际情况出发,实事求是,根据实际情况来做决策,找到我们应该怎么做。所以它是一个时代的产物,或者说是现实情况的一个反映,它并不是一个模仿的结果。 |
First of all, we don’t have anyone we’re modeling ourselves on. Every step we’ve taken has started from the actual situation on the ground — we make decisions realistically, based on facts, and figure out from there what we should do. So this is a product of the times, or a reflection of reality, rather than the result of imitating anyone. |
|
我们明确是要有商业化的。我们最终还是要能活下去,我们毕竟是一个公司,政府不会给我一分钱。 |
We are clearly committed to commercialization. Ultimately we still need to survive — we’re a company after all, and the government isn’t going to give us a single cent. |
|
我们本质上还是一个公司,只是说我们在考虑赚哪些钱、什么时候赚钱、赚多少钱、靠什么赚钱,我们有取舍。很多公司做得很伟大,因为它有一种利润以外的追求。那个追求最后不但没有影响到它的商业化,反而能让它商业化得更好。 |
At the end of the day, we’re still a company — it’s just that we’re thoughtful about which money to earn, when to earn it, how much to earn, and what to earn it from; we make trade-offs. Many great companies became great because they pursued something beyond profit. In the end, that pursuit didn’t hurt their commercialization at all — instead, it allowed them to commercialize even better. |
|
对于合作伙伴,其实我们这个融资是精心挑选的。首先我觉得,利益是比较一致的,就是跟我们利益最一致的、对我们最没有敌意的,或者说最希望我们能够做成功的。不是所有人都希望我们能做成功的,因为我们还是损害了很多其他人的利益的。 |
As for partners, we actually chose the investors in this financing round very carefully. First, I think the key is alignment of interests — those whose interests are most aligned with ours, who bear the least hostility toward us, or who most want to see us succeed. Not everyone actually wants us to succeed, because we have, after all, undercut the interests of many others. |
|
AI 现在不缺品位和直觉,它缺的是持续学习的能力。AI 的品位和直觉没有问题。你让它写个文章,它的品位和直觉,我觉得没有什么问题。 |
AI today isn’t lacking in taste or intuition — what it lacks is the ability to keep learning continuously. AI’s taste and intuition aren’t a problem. If you have it write an article, its taste and intuition, I think, are just fine. |
|
我们希望只做一块。我觉得 AI 这个事情很大,并不需要我……我只做一块。如果聚焦,并且我认为这里的生意利益已经足够大,就如果是 AI 时代会产生很多家万亿级别的公司,我觉得我们是其中一家。 |
We want to focus on just one piece of this. I think AI is such a huge undertaking that there’s no need for me to do everything — I’ll just do one piece. If we stay focused, and given that I believe the commercial opportunity here is already large enough — if the AI era is going to produce many trillion-dollar companies, I think we’ll be one of them. |
|
我们希望能够扶持更多的人,但是我们并没有那么多的精力。我们是有这个意愿,并且不会有利益冲突,但是我们有没有去做是另外一回事。但至少这里边是没有利益冲突的,我们是希望合作共赢的。 |
We would like to be able to support more people, but we don’t have the bandwidth to do that. We do have that intention, and there’s no conflict of interest involved — but whether we actually do it is a separate matter. At the very least, there’s no conflict of interest here; what we want is mutually beneficial cooperation. |


No comments:
Post a Comment