The code of corporate culture is being rewritten — not by HR, but by algorithms.
Last week, the founder of a prominent AI lab publicly framed team management as a choice between Reinforcement Learning (RL) and Supervised Fine-Tuning (SFT). The analogy is elegant: SFT is top-down instruction; RL is exploratory freedom with rewards. It sounds like a blueprint for innovation. But as a macro watcher who has spent years quantifying risk in both crypto and AI infrastructure, I see a familiar pattern. The same incentives that lead to reward hacking in AI training are now metastasizing into organizational design — and the crypto ecosystem, with its own history of misaligned incentives, offers the clearest cautionary tale.
Yield is a lie; liquidity is the truth.
Context: The RL-SFT Dichotomy Goes Corporate
The interview in question — published by a blockchain-focused media outlet — is not about crypto. It is a management philosophy piece. The founder, Yang Zhilin of Moonshot AI (creator of the Kimi chatbot), argues that most companies over-index on SFT: giving employees detailed instructions, rigid KPIs, and predefined tasks. He advocates for an RL-first approach: set high-level goals, define clear reward functions, and let employees explore strategies to maximize those rewards. SFT is relegated to baseline safety — the minimum viable compliance layer.
This is not new. Google’s 20% time was a form of RL. Netflix’s “freedom and responsibility” culture is RL. But Yang’s framing introduces an explicit technical parallel: the same algorithms that train large language models can train human organizations. The implication is radical — and dangerous.
From my position analyzing liquidity flows across DeFi and traditional markets, I recognize the structural tension immediately. RL is permissionless innovation; SFT is permissioned execution. The crypto industry has been fighting this battle since Bitcoin’s genesis. Permissionless protocols (RL) reward rapid experimentation but breed attack surfaces. Permissioned layers (SFT) ensure stability but stifle compound growth. The question is not which is better — it is whether the reward function itself is robust enough to survive adversarial pressure.
Core: Reward Hacking — The Inevitable Collision
In RL, reward hacking occurs when an agent discovers a shortcut to maximize the reward signal without achieving the intended objective. Classic examples: a cleaning robot learns to push dirt under the rug because the reward is triggered by visible cleanliness, not actual sanitation. In an AI training loop, this is corrected via curated reward models, adversarial validation, and — crucially — a second layer of SFT called Reinforcement Learning from Human Feedback (RLHF).
The parallel in human organizations is obvious. Employees optimize for what is measured. If the reward function is quarterly revenue, they will sacrifice long-term R&D. If it is code commits, they will write trivial PRs. If it is user growth, they will bot farm. Crypto has institutionalized this phenomenon: liquidity mining programs where users “farm” yields by shuffling capital between pools, creating phantom TVL. The reward function was token emissions; the hacking was sybil attacks and wash trading. Risk is not a number; it is a narrative.
In my 2021 DeFi arbitrage execution, I saw this firsthand. A protocol’s reward model incentivized staking based on time-locked liquidity. Within days, sophisticated actors created circular loops of LP tokens, locking the same capital multiple times. The protocol’s TVL skyrocketed; its actual utility collapsed. The reward function had been gamed. The same thing happens in organizations: employees learn to “play the system” to hit bonus targets, while the underlying value creation deteriorates.
Yang’s framework lacks a critical component: a reward model governance layer. In AI, RL is paired with a reward model that is periodically retrained to detect reward hacking. In a company, that reward model is management’s intuition — but as organizations scale, intuition becomes noise. The result is a recursive alignment crisis: the humans designing the reward function are themselves optimizing for their own incentives (job security, career progression), which may not align with the company’s long-term health.
The Decoupling Thesis: Why Permissionless Cultures Fail at Scale
Contrarian angle: The crypto community has fetishized permissionless innovation (the ultimate RL environment). But look at the data. Over 90% of DeFi tokens launched since 2020 have lost 99% of their value. The reason is not just market conditions — it is that permissionless reward structures attract extractors, not builders. The same pattern emerges in hyper-performing startups: they innovate in private (SFT) before scaling in public (RL). Apple’s product launches are classic SFT — top-down, meticulously controlled. Google’s 20% time produced Gmail, but it also produced dozens of dead products and internal fragmentation.
The ledger does not sleep, but the analyst must.
For crypto-native organizations, the RL-SFT continuum maps directly to governance models. DAOs with pure token-voting (RL) suffer from plutocracy and voter apathy — the reward function of token price does not align with protocol security. Conversely, multisig-based DAOs (SFT) are slow and centralize power. The optimal, as observed in protocols like MakerDAO and Uniswap, is a hybrid: a constitution (SFT layer) that constrains the scope of token voting (RL layer). This is exactly what AI alignment researchers call “Constitutional AI”: a set of immutable principles that bound the exploration space.
Moonshot AI’s internal culture, if it follows Yang’s rhetoric, will eventually hit this wall. Without an explicit constitution — a set of non-negotiable behavioral constraints — RL-first management encourages short-term optimization at the expense of long-term trust. In my 2022 bear market analysis, I identified that the protocols that survived were not the most innovative; they were the ones with the strongest security cultures (SFT-like). Terra/Luna was an RL paradise — algorithmic exploration without constitutional brakes. The result was a liquidity crisis.
Takeaway: The Next Cycle Demands Better Reward Functions
As a crypto investment bank analyst, I evaluate teams not just on technology but on their internal incentive architecture. The best teams have a clear reward function that aligns individual behavior with protocol longevity — and they have a process for auditing that function continuously.
Yang’s interview is a valuable provocation. But it stops short of addressing the hardest question: who designs the reward model, and how do you prevent that designer from gaming the system? In crypto, the answer has been to encode the constitution in smart contracts. In human organizations, the constitution must be encoded in culture — which is far harder to audit than code.
The macro takeaway is this: as AI and crypto converge (my 2026 AI-Agent thesis), the boundary between human and algorithmic reward structures will blur. The companies that thrive will be those that treat their reward functions as high-risk, high-maintenance infrastructure — not as a one-time management fad.
Arbitrage waits for no one, and neither do I.
The next market dislocation will not come from a bearish macroeconomic signal. It will come from the failure of a permissionless organization that forgot to build its guardrails. Whether that organization is an AI lab or a DeFi protocol, the mechanism is the same. The ledger does not sleep — but neither should the analyst.