{"version":"https://jsonfeed.org/version/1.1","title":"牧宇的Blog","home_page_url":"https://www.lihuanyu.com/","feed_url":"https://www.lihuanyu.com/feed.json","description":"以代码记行藏，以文字观万象；守寸心而求真，循长风以致远。","authors":[{"name":"skyADMIN","url":"https://www.lihuanyu.com"}],"items":[{"id":"https://www.lihuanyu.com/en/posts/2026/intelligence-is-becoming-aluminum/","url":"https://www.lihuanyu.com/en/posts/2026/intelligence-is-becoming-aluminum/","title":"Intelligence Is Becoming Aluminum","summary":"A personal reflection on shrinking frontend roles, increasingly capable coding agents, the falling price of intelligence, and why changing careers may be changing cabins rather than finding a lifeboat.","content_html":"<p>The company I work for has recently been reorganizing its development teams.</p>\n<p>What used to be a large frontend group is being split across individual business teams. Frontend hiring is shrinking. Roles are no longer defined as narrowly around frontend and backend specialties. The company increasingly wants one developer to follow a business need from interface to server and deliver the whole thing end to end.</p>\n<p>This can be described as an organizational adjustment. It can also be described as a move toward full stack. But for the people inside it, there is a more direct description: the same work is starting to require fewer people.</p>\n<p>I started using AI coding tools heavily in 2024. Cursor, Claude Code, Codex: I have used them through nearly every major jump in capability. Early AI felt like a fast typist. It was useful for filling in small functions, writing boilerplate, and explaining errors. Later, it could complete a module on its own. Now I can use natural language as the primary input and let AI handle most of the coding for an entire system.</p>\n<p>That is exciting.</p>\n<p>Projects that once needed several people now feel possible to start alone. The expansion of individual capability is real. But another question appears at the same time: if one person can work this way, why will a company still need so many programmers?</p>\n<p><a href=\"/posts/2026/%E6%99%BA%E5%8A%9B%E6%AD%A3%E5%9C%A8%E5%8F%98%E6%88%90%E9%93%9D/\">Chinese version of this article</a></p>\n<p><img src=\"/assets/posts/2026/intelligence-aluminum/01-intelligence-aluminum.jpg\" alt=\"A precious aluminum object opens into a mass-production line, representing intelligence changing from a scarce good into an industrial material\"></p>\n<h2>Frontend Is Only the First Alarm</h2>\n<p>Frontend roles currently seem to be taking the more visible hit. That does not mean AI is best at frontend work.</p>\n<p>In my own use, AI often writes ordinary server-side business code more smoothly than frontend pages. When requirements, data structures, and acceptance criteria are clear, server logic can be checked quickly through types, tests, and runtime behavior. Frontend work has to deal with rendering, visual detail, interaction states, browser compatibility, and real-device behavior. AI has to keep taking screenshots, comparing results, and trying again. A human often needs to watch more closely.</p>\n<p>Frontend may be shrinking first because its job boundary was the first to break. When one person with AI can cross old technical divisions, a company will naturally reorganize people around the business. For a developer, this is an expansion of capability. For the company, it is delivery with fewer people.</p>\n<p>Both descriptions are true. They are two sides of the same change.</p>\n<p><img src=\"/assets/posts/2026/intelligence-aluminum/02-one-developer-empty-desks.jpg\" alt=\"One developer handles frontend, server, and deployment work among a field of empty desks\"></p>\n<p>Moving from frontend to backend and becoming a full-stack developer is therefore a sensible choice today. It can help someone continue doing the work. It may not provide safety for much longer. Backend is not a shelter. Full stack is not a lifeboat.</p>\n<h2>Why Intelligence Looks Like Aluminum</h2>\n<p>Aluminum is abundant in the Earth’s crust, but it is difficult to isolate as a pure metal. In the early nineteenth century, extracting it was expensive, and aluminum briefly became a rare material used to display status. In 1886, the Hall-Heroult electrolytic process dramatically changed its production cost. Later industrial processes and cheap electricity helped turn aluminum from a precious object into an everyday material.</p>\n<p>Aluminum did not become useless. It appears in windows, beverage cans, power cables, cars, and aircraft. Humanity uses far more of it than it did when it was rare. Aluminum did not lose its value as a material. It lost the high price that came from scarcity.</p>\n<p>I increasingly think intelligence is going through a similar change. For now, I call it intelligence becoming aluminum.</p>\n<p>In the past, turning knowledge into code, a contract, a design, or an executable plan required a trained person to spend time. That person’s education, experience, and working hours were part of the price of the intellectual output. AI is turning this into a production process that can be invoked at scale. Models, compute, data, and electricity resemble a new set of electrolytic machinery.</p>\n<p>One reply is that generative AI is only a probability model, a recombination of human knowledge rather than genuine understanding or creativity. That argument can continue. Labor markets are usually less interested in philosophy. If AI-generated code runs, a proof survives verification, and a proposed solution solves the problem, the market will recalculate what it pays a human to produce the same result.</p>\n<p>Intellectual work that is describable, reproducible, and verifiable will certainly feel the pressure earlier. Code happens to satisfy all three conditions, which puts programmers near the front. I do not believe the pressure will stop there forever. As models begin handling harder problems in mathematics, science, and engineering, “creativity is the final defense” no longer sounds especially comforting.</p>\n<h2>More Software Does Not Necessarily Mean More Programmers</h2>\n<p>When aluminum became cheaper, the world produced more aluminum goods. When intelligence becomes cheaper, the world will produce more code, design, analysis, and content.</p>\n<p>This invites an easy mistake: if demand for software keeps growing, the number of programmers cannot fall. The amount of a product used and the number of people needed to produce it are not the same thing. Aluminum can spread across the world without raising the income of people who once produced it by hand.</p>\n<p>The software industry may not disappear. It may become more productive and more widespread than ever. But the programmer labor market built around large numbers of specialized roles, large teams, and long delivery cycles can still contract. The more efficient software production becomes, the fewer people each unit of demand needs. That puts real pressure on both jobs and wages.</p>\n<p>It may not arrive as one dramatic day when every programmer receives a layoff email. A company hires a few fewer people. An open role is not backfilled. Two teams become one. Work that used to be divided across several roles is repackaged into one job called “end-to-end ownership.” There is no ceremony. The number of people simply declines a little at a time.</p>\n<h2>Changing Cabins on the Titanic</h2>\n<p>Programmers are used to treating career anxiety as a learning problem. Frontend feels unsafe, so learn backend. Product engineering feels unsafe, so study architecture, algorithms, or AI. Coding feels unsafe, so move toward product, management, or consulting.</p>\n<p>Each move may help for a while. But all of them share one assumption: trained human intelligence remains scarce.</p>\n<p>If that assumption is weakening, moving from one knowledge job to another starts to look like changing cabins after the Titanic has hit the iceberg. Moving from third class to first class can buy comfort, access, and perhaps time. But once the ship is taking on water, an upgraded cabin is still not a lifeboat.</p>\n<p><img src=\"/assets/posts/2026/intelligence-aluminum/03-switching-cabins.jpg\" alt=\"Seawater enters an ocean liner while knowledge workers carry laptops between cabins\"></p>\n<p>What I see at work is less dramatic and more ordinary. When roles begin to disappear, the people who look most important are often the ones who know how to present work upward. I do not mean empty self-promotion alone. They can also explain a complicated problem clearly enough for an organization to make a decision. In real workplaces, those two abilities often live in the same person.</p>\n<p>AI can also write status reports, make slides, and summarize complex information. The mechanical act of producing a report will not be a durable advantage. What may remain scarcer is access to decision-makers, credibility inside the organization, and the power to define what counts as a result.</p>\n<p>That sounds less like a professional skill and more like a person’s position in an organization. Positions are hard to copy. They are also impossible for everyone to hold.</p>\n<h2>One Person Can Build a Product. That Does Not Make It a Business.</h2>\n<p>Independent development looks like another possible lifeboat.</p>\n<p>AI gives one person the productive capacity that once belonged to a small team. I also wonder whether individuals with AI can finally go toe to toe with established companies.</p>\n<p>I have built products. They have had users, and some have earned revenue. But the income has been badly out of proportion to the effort, and it is nowhere close to a salary. Code was not the real bottleneck. Distribution was. Without sustained promotion, there were not enough users. The only project that made noticeable money relied mainly on organic traffic from WeChat Mini Programs. What created revenue was not only my ability to build the product. It was also an existing distribution channel owned by WeChat.</p>\n<p>In China, a public-facing generative AI product also has to deal with filing requirements, content safety, and platform review. A large company can put legal, security, operations, and development teams around that list. An independent developer faces the whole list alone. Finishing the code only earns the right to enter the market.</p>\n<p>There is another complication. AI applications do not fully follow the low-marginal-cost logic of traditional software. Each additional successful use may trigger another model call, consuming more tokens and compute. Users arrive, and the bill arrives with them. I wrote about this separately in <a href=\"/en/posts/2026/ai-broke-the-zero-marginal-cost-myth-of-the-internet/\">Why AI Has Higher Marginal Costs Than Internet Software</a>.</p>\n<p>AI makes it easier for one person to build a product. It has not made it easier for one person to build a business.</p>\n<p><img src=\"/assets/posts/2026/intelligence-aluminum/04-product-distribution-gate.jpg\" alt=\"An independent developer produces a pile of applications but faces a narrow gate to the market, operating costs, and compliance paperwork\"></p>\n<p>As the development barrier falls, product supply grows and competition becomes more crowded. User attention, distribution, brand, trust, compliance capability, and the capital to carry continuing costs become more scarce. Unfortunately, many of those things still sit with large companies.</p>\n<p>Independent development is worth trying. It is not yet a proven lifeboat.</p>\n<h2>When Intelligence Gets Cheaper, What Gets More Expensive?</h2>\n<p>Every discussion about AI replacing a profession eventually arrives at the same comforting words: judgment, creativity, taste, responsibility, and human connection.</p>\n<p>Those abilities matter. I am less confident that they form a wall AI cannot cross. As generation, analysis, and experimentation become cheaper, the parts of those abilities that can be expressed and reproduced will also be repriced.</p>\n<p>What currently looks scarcer is decision-making power, customer relationships, distribution, organizational credibility, regulatory permission, and capital. Their common feature is that they are not only capabilities inside a person’s head. They are about which resources a person has the right to use, which relationships they can build, and how much of the output they can keep.</p>\n<p>That answer offers little comfort to an ordinary worker. If ownership of the intelligence factory becomes the valuable thing, understanding the argument does not suddenly give a salaried worker models, compute, distribution, and customers.</p>\n<p>Cheaper aluminum did not turn every aluminum fabricator into a shareholder of an aircraft company.</p>\n<h2>I Have Not Found the Lifeboat Yet</h2>\n<p>I still do not have a reliable answer.</p>\n<p>Learning full stack matters. It expands the work someone can do today. It does not prove that the future will need more full-stack engineers. Learning to present work matters. It helps good decisions become visible inside an organization. It does not preserve today’s number of roles forever. Building independent products matters. It lets someone face the market directly. It does not automatically produce users or income.</p>\n<p>None of that means doing nothing.</p>\n<p>Refusing AI is not an answer. It only makes someone lose today’s competitiveness sooner. Continuing to use AI, learning across the stack, building independent products, and practicing operations and communication are all worth doing. They simply should not be mistaken for a guarantee of career safety.</p>\n<p>AI has not delivered the leisure people imagined, either. Once I have paid for a subscription, unused tokens feel wasteful. Once AI finishes one thing faster, I immediately think of the next thing it could do. Time saved rarely becomes rest. More often, it becomes higher throughput. I become more capable, more tired, and more worried about the price of my own work at the same time.</p>\n<p>That is what I mean when I say intelligence is becoming aluminum.</p>\n<p>I do not know where the lifeboat is. At least I can stop mistaking a cabin upgrade for one. And I do not have to return to my cabin and sleep just because I cannot see the way out yet.</p>\n<p>Keep using the tools. Keep making things. Keep talking to users. Keep watching the waterline.</p>\n<p>None of this guarantees escape. But if a real lifeboat appears, it is better to already be on deck.</p>\n","date_published":"2026-08-09T00:00:00.000Z","tags":["AI","Programmers","Careers","Writing"],"language":"en"},{"id":"https://www.lihuanyu.com/posts/2026/%E6%99%BA%E5%8A%9B%E6%AD%A3%E5%9C%A8%E5%8F%98%E6%88%90%E9%93%9D/","url":"https://www.lihuanyu.com/posts/2026/%E6%99%BA%E5%8A%9B%E6%AD%A3%E5%9C%A8%E5%8F%98%E6%88%90%E9%93%9D/","title":"智力正在变成铝","summary":"从前端岗位收缩、AI 编程能力提升和独立开发经历出发，讨论智力成本下降后，程序员为何只是最早进水的舱室，以及在尚未找到救生艇时为什么仍要行动。","content_html":"<p>最近，我所在的公司正在重新组织开发团队。</p>\n<p>以前的大前端团队被拆到具体业务里，前端招聘开始收缩，岗位也不再只按前端、后端这些细分工种来定义。公司更希望一个开发者能跟着业务走，从界面到服务端，端到端完成一个需求。</p>\n<p>这套变化当然可以解释为组织调整，也可以解释为全栈化。但对身处其中的人来说，还有一个更直接的解释：同样的工作，开始不需要那么多人了。</p>\n<p>我从 2024 年开始密集使用 AI 编程工具。从 Cursor，到 Claude Code，再到 Codex，几乎每一轮能力变化都亲手试过。早期的 AI 更像一个打字很快的助手，适合补小方法、写样板代码、解释报错。后来它可以独立完成模块。到现在，我已经可以把自然语言作为主要输入，让 AI 完成一个系统的大部分编码工作。</p>\n<p>这件事让人兴奋。</p>\n<p>以前需要几个人配合的项目，现在一个人就敢开工。那种能力突然扩大的感觉很真实，但另一个问题也会跟着冒出来：既然一个人可以这样干，公司以后为什么还需要这么多程序员？</p>\n<p><a href=\"/en/posts/2026/intelligence-is-becoming-aluminum/\">English version: Intelligence Is Becoming Aluminum</a></p>\n<p><img src=\"/assets/posts/2026/intelligence-aluminum/01-intelligence-aluminum.jpg\" alt=\"珍贵铝器通向大规模铝生产线，象征智力从稀缺品变成工业品\"></p>\n<h2>前端只是先响的警报</h2>\n<p>前端岗位目前看起来受冲击更明显，但这不等于 AI 最擅长前端。</p>\n<p>我实际用下来，AI 写常规服务端业务代码，往往比写前端页面更顺。需求、数据结构和验收条件足够清楚时，服务端逻辑可以通过类型、测试和运行结果快速验证。前端反而要面对渲染效果、视觉细节、交互状态、浏览器兼容和真机表现。AI 得不断截图、比对、再修改，常常需要人盯得更紧。</p>\n<p>前端先收缩，可能只是因为它的岗位边界最先被打破。当一个人带着 AI 可以跨过过去的技术分工，公司自然会把人按业务重新组织。对开发者来说，这叫能力边界扩大；对公司来说，这叫用更少的人完成交付。</p>\n<p>两种说法都对，本来就是一件事的两面。</p>\n<p><img src=\"/assets/posts/2026/intelligence-aluminum/02-one-developer-empty-desks.jpg\" alt=\"一个开发者在空置工位之间同时处理前端、服务端和部署任务\"></p>\n<p>所以从前端转后端，再成为全栈开发者，当下当然是合理选择。它能让人继续完成工作，却未必能换来更长久的安全。后端不是避难所，全栈也不是救生艇。</p>\n<h2>智力为什么像铝</h2>\n<p>铝在地壳中并不缺，但它很难以单质形态被分离出来。19 世纪早期，铝的提炼昂贵，一度是适合展示身份的稀有金属。1886 年，霍尔-埃鲁法用电解大幅改变了铝的生产成本。再配合后来更完整的工业流程和便宜电力，铝从珍品变成了日常材料。</p>\n<p>今天的铝并没有变得没用。它出现在门窗、饮料罐、电缆、汽车和飞机里，用量远超它作为稀有金属的年代。铝的价值没有消失，消失的是它作为稀缺物的高价。</p>\n<p>我越来越觉得，智力也在经历类似的变化。我暂且把它叫作“智力铝化”。</p>\n<p>过去，要把知识变成一段代码、一份合同、一张设计图或一个可执行的方案，需要一个受过长期训练的人投入时间。这个人的学习成本、经验和工作时间，构成了智力产品的价格。AI 正在把这个过程改造成可以大规模调用的生产过程。模型、算力、数据和电力，就像一套新的电解设备。</p>\n<p>有人说生成式 AI 只是概率模型，是对人类知识的重组，不算真正的理解和创造。这个问题可以继续争论，但劳动市场通常没那么关心哲学。只要 AI 交付的代码能运行，给出的证明能通过验证，产生的方案能解决问题，市场就会重新计算人类做同样事情的价格。</p>\n<p>可描述、可复制、可验证的智力劳动，肯定会更早受到冲击。代码正好同时满足这三点，程序员因此站在最靠前的位置。但我不太相信冲击会永远停在这里。当模型开始处理越来越难的数学、科学和工程问题，“创造力是最后防线”也不再是一个让人放心的答案。</p>\n<h2>软件会更多，程序员未必</h2>\n<p>铝变便宜以后，铝制品变得更多。智力变便宜以后，代码、设计、分析和内容也会变得更多。</p>\n<p>这里容易出现一个误解：既然软件需求会继续增长，程序员就不会减少。可产品的用量和生产它的人数，从来不是一回事。铝的应用遍布世界，并不意味着手工提炼铝的人会获得更高收入。</p>\n<p>软件行业未必消失，甚至可能前所未有地繁荣。但过去那种依靠大量细分岗位、大量人力和长交付周期支撑起来的程序员行业，完全可能收缩。软件生产越高效，单位需求需要的人越少，岗位和工资所面对的压力就越真实。</p>\n<p>它不一定表现为某一天所有程序员同时收到裁员邮件。更可能的形态是：公司少招几个人，一个空缺不再补，两个团队合成一个，原本分给几个岗位的工作被重新包进一个“端到端负责”的职位里。没有哪一天锣鼓喧天，但人就这样慢慢少了。</p>\n<h2>泰坦尼克号上的换舱</h2>\n<p>程序员习惯用学习解决职业焦虑。前端不安全，那就学后端；业务开发不安全，那就学架构、算法或 AI；写代码不安全，那就往产品、管理或咨询走。</p>\n<p>这些选择都可能在某个阶段有用，但它们共享同一个前提：经过长期训练的人类智力仍然是稀缺品。</p>\n<p>如果这个前提正在动摇，那么从一个脑力岗位转向另一个脑力岗位，就像泰坦尼克号撞上冰山以后换舱室。从三等舱换到头等舱，短期体验当然不同；船真的开始下沉时，升舱本身不是逃生方案。</p>\n<p><img src=\"/assets/posts/2026/intelligence-aluminum/03-switching-cabins.jpg\" alt=\"海水进入船舱后，仍有人带着电脑在不同舱室之间移动\"></p>\n<p>我在公司里看到的情况更现实一点。当岗位开始减少，显得更重要的人往往很会汇报。这里的“会汇报”不只是包装成果、向上管理，也包括把复杂问题讲清楚，让组织可以作出决策。两种能力在真实环境里经常长在同一个人身上。</p>\n<p>AI 也能写周报、做 PPT、总结复杂信息。只会写汇报材料，不会成为长期壁垒。比表达更稀缺的，可能是接近决策者的机会、被组织承认的信用，以及定义什么算成果的权力。</p>\n<p>这听起来不像职业技能，更像一个人在组织里的位置。位置不容易被复制，可位置本来就不可能人人都有。</p>\n<h2>一个人做出产品，不等于做成生意</h2>\n<p>另一个看起来像救生艇的方向，是独立开发。</p>\n<p>AI 让一个人具备过去一个小团队的生产能力。我也会想，个体有了 AI 以后，是不是终于有机会和传统大公司掰掰手腕。</p>\n<p>我确实做出过产品，也确实有过用户和收入。但收入和投入完全不成比例，更不能和工资相比。真正卡住我的也不是代码，而是运营。没有持续宣传，就没有足够用户。唯一明显赚到钱的项目，主要依靠微信小程序的自然流量。带来收入的不只是把东西做出来的能力，还有微信已经存在的分发渠道。</p>\n<p>在国内做面向公众的生成式 AI 产品，还要面对备案、内容安全和平台审核等现实门槛。一个大公司可以让法务、安全、运营和开发一起处理，个人开发者面对的却是一整张清单。写完代码，只是拿到了进入市场的资格。</p>\n<p>更麻烦的是，AI 应用不完全符合传统软件的低边际成本逻辑。每多一次有效使用，背后都可能多一次模型调用、多一份 token 和算力成本。用户来了，账单也来了。我在《<a href=\"/posts/2026/AI%E6%89%93%E7%A0%B4%E4%BA%86%E4%BA%92%E8%81%94%E7%BD%91%E7%9A%84%E9%9B%B6%E8%BE%B9%E9%99%85%E6%88%90%E6%9C%AC%E7%A5%9E%E8%AF%9D/\">AI 打破了互联网的零边际成本神话</a>》里单独写过这个问题。</p>\n<p>所以，AI 让一个人更容易做出产品，却没有让一个人更容易做成生意。</p>\n<p><img src=\"/assets/posts/2026/intelligence-aluminum/04-product-distribution-gate.jpg\" alt=\"独立开发者生产出大量应用，却被分发门槛、成本和合规挡在市场之外\"></p>\n<p>开发门槛下降后，产品供给会更多，竞争会更挤。用户注意力、渠道、品牌、信任、合规能力和承担持续成本的资本，反而更显得稀缺。不巧的是，这些东西很多仍然掌握在大公司手里。</p>\n<p>独立开发值得尝试，但它还不是一条经过验证的救生艇。</p>\n<h2>智力降价后，什么会更贵</h2>\n<p>每当讨论 AI 会不会替代某个职业，最后总会剩下几个让人安心的词：判断力、创造力、品味、责任感，还有人与人的连接。</p>\n<p>这些能力当然重要。但把它们直接当成 AI 无法穿透的护城河，我还没有这么乐观。当生成、分析和试错的成本继续下降，这些能力里可以表达和复制的部分，同样会被重新定价。</p>\n<p>目前看起来更稀缺的，是决策权、用户关系、分发渠道、组织信用、合规资格和资本。它们的共同点是，不只是一个人脑子里的能力，而是一个人有权调用什么资源，能和谁建立关系，又能拿走多少产出。</p>\n<p>可这个答案也不太能安慰普通人。如果未来最贵的是对智力工厂的所有权，那么一个依靠工资生活的人，不会因为听懂了这个道理，就突然拥有模型、算力、渠道和客户。</p>\n<p>铝便宜了，并不意味着每一个做铝制品的人都变成飞机公司的股东。</p>\n<h2>还没有找到救生艇</h2>\n<p>到现在，我还没有找到一个可靠答案。</p>\n<p>学全栈很重要，它可以扩大当下的工作范围，却不能证明未来需要更多全栈工程师。学会汇报很重要，它能让正确的事被组织看见，却不会让组织永远保留现在的岗位数量。做独立产品很重要，它让人尝试直接面对市场，却不会自动带来用户和收入。</p>\n<p>可这不等于什么都不用做。</p>\n<p>拒绝 AI 不是办法。那只会让人更早失去当下的竞争力。继续使用 AI、学习全栈、尝试独立产品、练习运营和表达，也都值得做。只是做这些事时，不再假设它们一定能换来职业安全。</p>\n<p>AI 也没有带来想象中的轻松。订阅买了，token 没用完，会觉得浪费；AI 把一件事做快了，人就会立刻想做下一件。节省出来的时间很少变成休息，往往只是让工作的吞吐量继续上升。个人变得更强，也变得更累，同时还要担心自己的价格。</p>\n<p>这就是我对“智力铝化”最直接的感受。</p>\n<p>我不知道救生艇在哪里。但至少可以不再把升舱当成救生艇，也不因为还没看见出口，就回舱里继续睡觉。</p>\n<p>继续用工具，继续做东西，继续接触用户，也继续看水线到了哪里。</p>\n<p>这些事不保证能逃出去。但真的救生艇出现时，人最好已经在甲板上。</p>\n","date_published":"2026-08-09T00:00:00.000Z","tags":["AI","程序员","职业","随笔"],"language":"zh"},{"id":"https://www.lihuanyu.com/en/posts/context-engineering-is-not-about-more-context/","url":"https://www.lihuanyu.com/en/posts/context-engineering-is-not-about-more-context/","title":"Context Engineering Is Not About Adding More Context","summary":"A large context window only defines how much a model can see. Context engineering decides what enters that window, how to compress and order it, what to isolate, and what to discard as a task grows.","content_html":"<p>An LLM feature rarely collapses under context on its first day.</p>\n<p>Version one is lean: a system prompt, a user request, and one model call. After the first wrong answer, the prompt gains another rule. The application then adds retrieval for private documents, tools for live data, conversation history for continuity, and task logs for recovery.</p>\n<p>Each addition makes sense on its own. A request captured a few months later may contain instructions, documents, search results, tool output, chat history, user preferences, failure records, and an old summary. The model now sits at a desk covered in paper. Everything is within reach, but the page that matters is harder to find.</p>\n<p>Saying “the model has a large context window” does not solve that problem.</p>\n<p><img src=\"/assets/posts/2026/context-engineering-en/01-context-assembly-pipeline.jpg\" alt=\"Documents pass through selection, compression, ordering, isolation, and deletion before entering a bounded working context\"></p>\n<p><a href=\"/posts/context-engineering-is-not-more-context/\">Chinese version of this article</a></p>\n<p><a href=\"/en/posts/2026/from-one-model-call-to-an-agent/\">From One Model Call to an Agent</a> maps prompts, context, retrieval-augmented generation (RAG), memory, tools, and agents. This article takes one part of that map further. It examines what an application should put into a model call and what it should leave outside.</p>\n<h2>A larger window gives capacity, not quality</h2>\n<p>A context window sets the maximum number of tokens available to one inference. It works like the surface area of a desk. A larger desk holds more material, but it does not organize that material or make the right page easier to notice.</p>\n<p>Long context creates four distinct problems.</p>\n<p>The first is cost and latency. Longer input increases transfer size and the model’s prefill work before it generates the first output token. In an agent loop, later calls may resend material from every earlier step, so one extra block of context can incur cost several times.</p>\n<p>The second is interference. If only two of ten retrieved documents support the answer, the other eight are not free background. They compete with the useful evidence. Stale rules, duplicate summaries, and conflicting records make the problem worse because the model must infer which source deserves trust.</p>\n<p>The third is position. Research such as <a href=\"https://arxiv.org/abs/2307.03172\">Lost in the Middle</a> shows that models can use information unevenly across a long input. A fact’s presence in the context does not guarantee reliable use. Burying a critical constraint between thousands of tokens and ending with “follow all instructions above” expresses hope, not control.</p>\n<p>The fourth problem is output space. The context window must also leave room for the answer. When input consumes the budget, the model cannot finish a report, code patch, or JSON object. Some APIs reject the call, while some runtimes truncate earlier messages. Neither outcome belongs in a production strategy.</p>\n<p>A larger window solves “this material does not fit.” Context engineering addresses different questions: what belongs, where it belongs, and when it should leave.</p>\n<h2>What occupies one model call</h2>\n<p>Context includes more than the latest user message. A tool-using application may assemble all of these sources for one call:</p>\n<ul>\n<li>System instructions and safety rules</li>\n<li>The current task and the user’s latest constraints</li>\n<li>Evidence retrieved from documents, databases, or search</li>\n<li>Results from recent tool calls</li>\n<li>Recent conversation and task state</li>\n<li>A small set of long-term memories</li>\n<li>An output schema or format example</li>\n</ul>\n<p>The application must reserve space for the model’s output as well.</p>\n<p><img src=\"/assets/posts/2026/context-engineering-en/02-context-budget.jpg\" alt=\"A context window allocates separate areas for instructions, the current task, evidence, tool results, recent state, and output\"></p>\n<p>These sources do not have equal status. Instructions stay stable. The current task must remain intact. Retrieval evidence can compete on relevance, while old tool logs may only need a compact conclusion. Conversation history loses value as the task changes.</p>\n<p>Concatenating everything into one string works until the budget runs out. It also hides which source caused the failure. Separate sections make it possible to assign each source a budget and an eviction policy.</p>\n<p>A single global rule such as “drop the oldest message when over budget” is too blunt. The oldest message may contain the task goal. I prefer to reserve output space first, protect instructions and the current task, let evidence and tool results share a flexible area, and retain only the state needed to continue.</p>\n<p>That policy is conservative. It is still better than receiving half a JSON object because the input took the model’s last available tokens.</p>\n<h2>Context engineering has five operations</h2>\n<p>Context engineering extends beyond prompt wording. Prompt engineering asks how to express the task. Context engineering decides where supporting material comes from, how it enters a call, and how it leaves as the task grows.</p>\n<p>The implementation usually contains five operations.</p>\n<p><strong>Selection.</strong> Find candidate material for the current question, then choose what enters the working set. A shipping request does not need the full refund policy. A database timeout investigation does not need every frontend log.</p>\n<p><strong>Compression.</strong> Reduce long source material while preserving facts, provenance, and time. A graceful paragraph that loses the error code, amount, version, or file path is not a useful summary. The missing detail may be impossible to recover later.</p>\n<p><strong>Ordering.</strong> Keep stable rules in a stable location, mark the current task, and place important evidence near the instruction that uses it. When sources differ in authority or age, label those differences instead of asking the model to guess.</p>\n<p><strong>Isolation.</strong> Give each subtask its own working set. A retrieval step may not need writing-style rules. The model composing the final answer does not need raw tool debugging logs. Isolation also reduces the amount of sensitive data exposed to each call.</p>\n<p><strong>Deletion.</strong> Remove expired plans, duplicate observations, and summaries superseded by newer facts. Deletion is part of normal long-task execution, not an emergency response. Memory that only accepts writes eventually becomes another log archive.</p>\n<p>None of these operations is mysterious. The hard part is that they cannot live in a prompt template alone. Retrieval, state storage, budget calculation, and the runtime must enforce them together.</p>\n<h2>RAG, memory, and tool results need different policies</h2>\n<p>Several context sources look similar after they become text. They still require different retention rules.</p>\n<div class=\"table-scroll\"><table>\n<thead>\n<tr>\n<th>Source</th>\n<th>What it should contain</th>\n<th>How it enters context</th>\n<th>When it leaves</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Prompt / instructions</td>\n<td>Goal, rules, output format</td>\n<td>Retained in a stable section</td>\n<td>Task completion or rule update</td>\n</tr>\n<tr>\n<td>RAG</td>\n<td>External facts relevant to the current question</td>\n<td>Retrieved, ranked, deduplicated, and labeled</td>\n<td>Query change, staleness, or low relevance</td>\n</tr>\n<tr>\n<td>Memory</td>\n<td>Preferences, durable facts, prior conclusions</td>\n<td>Retrieved first, then selected in small amounts</td>\n<td>Update, low confidence, or lost relevance</td>\n</tr>\n<tr>\n<td>Tool result</td>\n<td>An observation from the current step</td>\n<td>Key raw fields stay recent; older results get compacted</td>\n<td>Extracted conclusion or newer result</td>\n</tr>\n<tr>\n<td>Runtime state</td>\n<td>Stage, step count, failures, pending work</td>\n<td>Encoded in a compact structure</td>\n<td>Task completion or stage transition</td>\n</tr>\n</tbody>\n</table>\n</div><p>RAG involves more than returning the top few text chunks. The system must decide what to search, how many results to keep, how to remove duplicates, and how to preserve source and freshness metadata. Memory has a similar distinction: storing a fact does not make the model know it. The application must retrieve that fact when it becomes relevant.</p>\n<p>Tool results create the fastest context growth. A log tool may return hundreds of lines. The model reads them and calls another tool, while the runtime appends both raw results and both explanations to the conversation.</p>\n<p>After three or four steps, the original task may occupy a small corner of the input. A safer policy keeps the latest observations in detail and compresses older results into a checkpoint with evidence pointers.</p>\n<h2>Long tasks get hard on the second call</h2>\n<p>You can assemble context for one call by hand. The second call introduces the real design work: how much of the first answer stays, whether tool output returns verbatim, whether failed attempts matter, and how new evidence overrides old conclusions.</p>\n<p>Assume each step adds an observation of similar size and every later call resends the full history. Call ten pays for the first ten observations, not only the tenth. Across the task, cumulative input grows approximately with the square of the step count.</p>\n<p>An unbounded loop therefore expands the context window, latency, and bill together. A long task needs a context lifecycle, not a chat transcript that grows forever.</p>\n<p><img src=\"/assets/posts/2026/context-engineering-en/03-long-task-context-lifecycle.jpg\" alt=\"A long-running task retrieves, acts, observes, compacts, checkpoints, and resumes without resending its full transcript\"></p>\n<p>The most recent steps can retain detail. Older steps become a checkpoint. A useful checkpoint records the goal, confirmed facts, completed actions, failed attempts, and missing information.</p>\n<p>That record must support recovery after an interruption. “The task made good progress” cannot restart anything.</p>\n<p>Once a subtask finishes, its process material can leave the parent context. Only its conclusion, evidence, and unresolved questions return. A later stage retrieves source material again when it needs the details.</p>\n<p>This approach performs more assembly work between calls. In return, each model invocation receives a bounded workspace with a clear purpose.</p>\n<h2>A practical context assembly step</h2>\n<p>You do not need a large agent framework to establish these boundaries. Separate material collection from the model call first.</p>\n<pre><code class=\"language-typescript\">const candidates = await collectCandidates(task, state)\nconst evidence = selectAndRank(candidates, task)\nconst recent = keepRecentObservations(state.events, 2)\nconst checkpoint = compactOlderEvents(state.events)\n\nconst context = assemble({\n  instructions: stableInstructions,\n  task: normalizeTask(task),\n  evidence: fitToBudget(evidence, budgets.evidence),\n  recent,\n  checkpoint,\n  outputSchema,\n  reserveForOutput: budgets.output,\n})\n\nconst result = await model.generate(context)\nawait persistResultAndState(result, state)\n</code></pre>\n<p>The important part is the boundary between each function. Candidate material is not the final context. Recent observations and older state use different policies. Output space affects assembly before the call, and the runtime persists state outside the context after the call.</p>\n<p>A mature implementation can attach metadata to each block: <code>source</code>, <code>timestamp</code>, <code>priority</code>, <code>ttl</code>, and <code>sensitivity</code>. Provenance, expiry, and permission should not exist only as prose inside the material.</p>\n<h2>Measuring whether context engineering works</h2>\n<p>Lower token usage is one metric, not the result. Removing too much context can produce a faster wrong answer.</p>\n<p>I track three groups of measurements.</p>\n<p>The first covers cost and latency: input tokens per call, cumulative tokens per task, time to first token, total duration, and the p50 and p95 distributions. An average can hide a small set of requests that carry most of the context.</p>\n<p>The second covers task quality. Use a fixed set of real requests and verify the evidence behind each answer. For structured output, run schema and domain validation. For tools, inspect the retrieval path as well as the final prose.</p>\n<p>The third covers runtime behavior: candidate count, selected count, truncation by section, checkpoint reuse, and stop reason. Without these records, context failures collapse into “the model is unstable,” which does not identify a fix.</p>\n<p>An evaluation set does not need hundreds of cases at the start. Twenty or thirty tasks that previously failed can expose useful differences. Fix the input, required evidence, and pass condition before comparing a new model or assembly policy.</p>\n<h2>Failure patterns worth watching</h2>\n<p>The most common mistake is treating “the model may need this later” as “include this on every call.” Potentially useful material belongs in retrievable storage. It enters context when the current step needs it.</p>\n<p>Premature summarization creates another failure. The system saves tokens by replacing source text with a smooth paragraph, then discovers that the paragraph omitted the number that decides the case. Keep source pointers in summaries and retain original excerpts for high-risk facts.</p>\n<p>Applications also give every source equal trust. A live database query, a three-month-old chat summary, a new user statement, and the model’s previous guess should not share one level. Record provenance, freshness, and confidence explicitly.</p>\n<p>Finally, some runtimes wait for an overflow error and then remove the oldest messages. The API returns <code>200</code>, but the task goal may have disappeared. A controlled degradation policy works by section: deduplicate, drop low-relevance evidence, compact older observations, and protect the task and stable rules until the end.</p>\n<h2>The Novevia implementation</h2>\n<p>I met this problem in Novevia, an AI fiction-writing application. Its first whole-book review feature compressed the setting, chapter cards, facts, and memories into one large request. As a book grew, the first request grew with it. Two model calls once took more than eight minutes.</p>\n<p>I stopped trying to write a stronger giant prompt and changed context assembly instead. The first call now receives a manifest. The model retrieves chapters, story routes, or facts when it needs them. Each call has a separate budget, older tool results get compacted, and task state remains in the database until a later step needs it.</p>\n<p>The design also forms a bounded ReAct-style loop, but ReAct did not remove the timeout by itself. The smaller working set, timed retrieval, context eviction, and stop conditions did the work.</p>\n<p><a href=\"/en/posts/2026/i-stopped-sending-the-whole-book-to-the-model/\">I Stopped Sending the Whole Book to the Model</a> describes the read-only tools, step limits, validation, recovery, and rollback behind that implementation. Novevia supplies one concrete case. The assembly problem appears anywhere an LLM task lasts beyond one call.</p>\n<p>Too little context forces the model to guess. Moving the entire archive into the window creates a different failure.</p>\n<p>A practical target is narrower: give each call the material required for its current step, without enough unrelated material to obstruct the work. Context engineering starts there.</p>\n","date_published":"2026-08-08T00:00:00.000Z","tags":["AI","LLM","Context Engineering","RAG","Agent"],"language":"en"},{"id":"https://www.lihuanyu.com/en/posts/2026/from-one-model-call-to-an-agent/","url":"https://www.lihuanyu.com/en/posts/2026/from-one-model-call-to-an-agent/","title":"From One Model Call to an Agent: How LLMs, RAG, Tools, Loops, and ReAct Fit Together","summary":"LLMs, prompts, context, RAG, tools, workflows, loops, ReAct, memory, and agents belong to different layers. This guide explains what each one does and how they fit into a working system.","content_html":"<p>Connect an application to a model API and it can chat. Add a search endpoint and the product description starts saying RAG. Give the model a few tools and the word agent appears. Put a <code>while</code> loop around it, and ReAct sometimes joins the list.</p>\n<p>The terminology grows faster than the code.</p>\n<p>The trouble is that these terms do not describe the same layer. A large language model (LLM) is the model. Prompt and context are inputs. Retrieval-augmented generation (RAG) and tools add external information or capabilities. Workflows, loops, and ReAct organize execution. An agent is the running system that combines those parts. Protocols and frameworks such as the Model Context Protocol (MCP) and Pi sit farther outside.</p>\n<p>Four questions cut through most of the naming: what can the model see on each call, who chooses the next step, who executes external actions, and what makes the task stop?</p>\n<p>Start there and the map becomes much less mysterious.</p>\n<p><img src=\"/assets/posts/2026/llm-agent-concepts-en/01-llm-agent-concept-map.jpg\" alt=\"A layered map of LLMs, context, external capabilities, orchestration, state, and agent boundaries\"></p>\n<p><a href=\"/posts/from-one-model-call-to-an-agent/\">Chinese version of this article</a></p>\n<h2>An LLM is the model, not the application</h2>\n<p>A large language model is the generation and decision component inside an AI application.</p>\n<p>At the training level, it predicts later tokens from earlier ones. At the application level, it receives messages and produces text or structured output. The act of running the model to produce that output is inference.</p>\n<p>The smallest useful call may contain one instruction:</p>\n<pre><code class=\"language-text\">Summarize the following passage in three sentences.\n</code></pre>\n<p>The model receives the instruction and source text, returns an answer, and the call ends. This is an ordinary LLM call. It needs no tool, loop, or agent.</p>\n<p>The model is not the rest of the application. The interface, user identity, database, permissions, conversation history, retries, and billing all belong to surrounding systems. A chat interface may appear to remember earlier messages because the application or model service stores them and supplies relevant parts on later calls.</p>\n<p>From the model’s point of view, the current output depends on the current context. Its parameters do not acquire a private memory of one user because it answered that user yesterday.</p>\n<h2>Prompt defines the task; context defines what the model can see</h2>\n<p>Developers use prompt and context as if they mean the same thing. Their responsibilities are different.</p>\n<p>A prompt describes the task: the role, goal, rules, and output format. System instructions, the user’s request, and format constraints belong here.</p>\n<p>APIs do not use the word prompt consistently. Some use it for the entire model input. I use a narrower definition in this article because it makes the engineering responsibilities easier to see: prompt expresses the task, while context contains everything available to the model on that call.</p>\n<p>Context may include:</p>\n<ul>\n<li>System instructions and the user’s request</li>\n<li>Source material to process</li>\n<li>Conversation history</li>\n<li>Retrieved documents</li>\n<li>Tool definitions</li>\n<li>Results returned by earlier tool calls</li>\n</ul>\n<p>The context window sets the maximum number of tokens a call can contain. A token is the model’s basic text unit. It may be one Chinese character, part of an English word, or punctuation, so token count and character count are not interchangeable.</p>\n<p>Prompt engineering asks how to express the instruction. Context engineering also asks where the material comes from, when to include it, how long to retain it, and what to discard when the budget runs out.</p>\n<p>Think of an LLM as someone working at a desk. The prompt is the assignment. Context is every document on the desk. The context window is the desk’s surface area. A larger desk holds more paper, but emptying the archive onto it does not improve judgment.</p>\n<h2>RAG adds external evidence before generation</h2>\n<p>Once trained, a model does not automatically know a company’s private policies, breaking news, or the latest row in a database. <a href=\"https://arxiv.org/abs/2005.11401\">Retrieval-augmented generation</a> adds that information at inference time.</p>\n<p>Consider a question about a company’s travel expense limit. The model can guess from its training data, or the application can retrieve the relevant policy and ask the model to answer from that evidence.</p>\n<p>The second path looks like this:</p>\n<pre><code class=\"language-text\">Question -&gt; retrieve relevant material -&gt; add it to context\n         -&gt; call the model -&gt; generate the answer\n</code></pre>\n<p>RAG changes the material supplied to one inference. It does not change model parameters. When the policy changes, the application updates the document store instead of retraining the model.</p>\n<p>That distinction also separates RAG from fine-tuning. RAG supplies current facts during inference. Fine-tuning changes model parameters and suits stable behavior, style, or task-specific capability. An application can use both, but they solve different problems.</p>\n<p>RAG is not another name for a vector database. Embeddings turn content into vectors that support semantic comparison, and vector databases store and search those vectors. They are a common retrieval implementation, not the definition of RAG.</p>\n<p>Full-text search, keyword matching, SQL queries, and direct file reads can all retrieve material for generation. If the system fetches relevant information and uses it to produce the answer, the basic retrieval-augmented structure is present.</p>\n<p>A fixed RAG pipeline may still be an ordinary workflow. Server code can create a query, fetch the top results, and call the model once. It needs neither an agent nor ReAct.</p>\n<h2>Tools let the model request external capabilities</h2>\n<p>RAG mainly addresses missing information. Tools cover both reading and action.</p>\n<p>Order lookup, log search, and file reads are read operations. Sending an email, creating a ticket, changing a calendar event, and starting a deployment are write operations. The model usually owns none of those permissions. It can only express an intent to call a tool.</p>\n<p>With function calling, the application tells the model each tool’s name, purpose, and parameter schema. A shipping question may produce output like this:</p>\n<pre><code class=\"language-json\">{\n  &quot;tool&quot;: &quot;get_order&quot;,\n  &quot;args&quot;: { &quot;orderId&quot;: &quot;1234567890123&quot; }\n}\n</code></pre>\n<p>The JSON does not query a database by itself. The runtime must validate the order ID and the current user’s permissions, execute <code>get_order</code>, and return the result to the model for the final response.</p>\n<p>The model proposes what it wants to call. The runtime decides whether the call is allowed and how to execute it. Without that division, any text that resembles a command could become an instruction to the production system. That arrangement eventually produces an incident.</p>\n<p>Tools and RAG can overlap. If knowledge search is exposed as a tool and its result enters the generation process, the operation is both tool use and RAG. Sending an email is tool use, but it is not usually called RAG because it changes an external system instead of retrieving evidence for an answer.</p>\n<p>One tool call does not automatically create an agent. A fixed sequence that looks up one order and writes one response is still a workflow.</p>\n<h2>Workflow, loop, and ReAct answer three different questions</h2>\n<p>Comparing workflow, loop, and ReAct as alternatives creates confusion. They describe three different properties of execution.</p>\n<p><img src=\"/assets/posts/2026/llm-agent-concepts-en/02-workflow-loop-react.jpg\" alt=\"Workflow presets a path, a loop repeats under a condition, and ReAct chooses the next step from observations\"></p>\n<p>A workflow asks who arranges the steps. A production incident report might always check service status, read recent errors, collect deployment records, and ask the model to write a report. Code has already fixed the path. The model only processes material at selected points.</p>\n<p>A loop asks whether the process repeats. Three retries form a loop. Polling a task until completion forms a loop. Asking a model to revise code until tests pass also forms a loop. Repetition is a control structure, not intelligence.</p>\n<p><a href=\"https://arxiv.org/abs/2210.03629\">ReAct</a> asks how the model chooses the next step inside a loop. The name comes from Reasoning and Acting. The model evaluates current information, selects an Action, receives an Observation, and then evaluates the new state.</p>\n<p>An incident investigation might follow this route:</p>\n<pre><code class=\"language-text\">Observe the alert\n-&gt; inspect error logs\n-&gt; find database connection timeouts\n-&gt; inspect connection pool settings\n-&gt; find a recent configuration change\n-&gt; inspect deployment records\n-&gt; produce a diagnosis\n</code></pre>\n<p>The outer runtime continues or stops the loop. The model selects each Action from the latest Observation. Server code does not prescribe the query order in advance; the evidence changes the path.</p>\n<p>A loop is not the same thing as ReAct. A fixed workflow can run inside a loop, and a retry loop needs no model at all. ReAct is one model-driven decision pattern inside a loop. It is not the only way to build an agent.</p>\n<p>ReAct also does not require an application to store a long chain of thought. A short action reason, tool arguments, and observations are enough to reconstruct the operational path. Auditing should focus on what the model requested, what the system returned, and which rule allowed execution to continue.</p>\n<h2>An agent combines the model, tools, state, and runtime</h2>\n<p>Agent has no boundary that every vendor or developer accepts. Some products add the word as soon as a model calls one tool. An implementation still needs a definition that maps to code.</p>\n<p>I find this one useful: an agent is a running system that keeps making decisions and taking actions toward a goal. A production agent usually contains these parts:</p>\n<ul>\n<li><strong>Model</strong>: interprets the current state and chooses the next step</li>\n<li><strong>Instructions and context</strong>: supply the goal, rules, and current material</li>\n<li><strong>Tools</strong>: read information or change external systems</li>\n<li><strong>State</strong>: records task progress and prior events</li>\n<li><strong>Control loop</strong>: moves between model decisions, Actions, and Observations</li>\n<li><strong>Stop conditions</strong>: define completion, failure, timeout, and human handoff</li>\n<li><strong>Boundaries</strong>: limit permissions, budget, steps, and dangerous operations</li>\n</ul>\n<p>An imprecise but memorable formula is:</p>\n<pre><code class=\"language-text\">Agent = Model + Context + Tools + State + Loop + Boundaries\n</code></pre>\n<p><img src=\"/assets/posts/2026/llm-agent-concepts-en/03-bounded-agent-system.jpg\" alt=\"An agent runtime with a model, tools, state, stop conditions, and permission boundaries\"></p>\n<p>For a production incident, an agent may read metrics, logs, configuration, and deployment history. It can preserve clues across several queries and submit a diagnosis. Code still limits the total time, maximum steps, and available tools. A write operation can pause for human approval.</p>\n<p>Maximum autonomy is not the goal. Production systems need read-only tools, call budgets, timeouts, approvals, validation, idempotency, and rollback. None of this is glamorous. It decides whether a failed agent leaves behind a failed task or a damaged system.</p>\n<p>Workflow and agent are not binary categories. The more choices code fixes in advance, the closer the system is to a workflow. The more paths the model can choose based on current state, the more agentic it becomes. Reliable systems combine both: the model explores, while fixed code validates, approves, and executes.</p>\n<h2>Memory is not a larger context window</h2>\n<p>An agent that works across several calls or tasks needs to preserve some information outside the model invocation.</p>\n<p>Context is what the model can see now. Memory is information the application stores and may retrieve later. It can live in a database, file, or dedicated memory service. Saving it does not make the model aware of it. The application must retrieve the relevant part and place it in the current context.</p>\n<p>Application memory usually includes at least three categories:</p>\n<ul>\n<li><strong>Conversation state</strong>: what has been said in the current conversation</li>\n<li><strong>Domain memory</strong>: user preferences, business facts, document summaries, and other durable information</li>\n<li><strong>Runtime state</strong>: current task position, completed tool calls, and failure count</li>\n</ul>\n<p>Keeping every historical message in every later call is not memory design. It is an ever-growing context. Memory design must decide what to store, when to retrieve it, which source wins when records conflict, and how stale information expires.</p>\n<h2>MCP, Pi, and agent frameworks are infrastructure</h2>\n<p>Once the model, input, tools, and orchestration are separated, protocols and frameworks become easier to place.</p>\n<p>The <a href=\"https://modelcontextprotocol.io/docs/getting-started/intro\">Model Context Protocol</a> defines a standard way for a host to connect to external tools, resources, and prompts. It answers how capabilities are exposed and connected. It does not choose the agent’s goal, loop strategy, or permission boundary.</p>\n<p>Projects such as Pi, LangChain, and Mastra handle some shared engineering work. They may normalize model providers, organize messages, declare tools, run loops, record traces, or connect to external services. Their scope differs, so the label Agent Framework does not say enough by itself.</p>\n<p>Pi here means the libraries around <code>@earendil-works/pi-ai</code>, not Raspberry Pi. In one project, I evaluated <code>pi-ai</code> as a model integration layer. It normalizes providers, messages, streaming, tool calls, and usage. A related agent runtime can remove part of the hand-written loop and tool-execution code.</p>\n<p>The domain rules remain in the application: which data a tool may read, which actions require approval, and what result counts as complete. A framework supplies reusable runtime pieces. MCP standardizes connections. Neither one decides the product’s policy.</p>\n<h2>Put the concepts on one map</h2>\n<p>The concepts fit into a layered view of an LLM application:</p>\n<pre><code class=\"language-text\">Application goal and boundaries\n└── Agent\n    ├── Orchestration: Workflow / Loop / ReAct\n    ├── External capabilities: RAG / Tool\n    ├── State: Memory / Runtime State\n    ├── Current input: Prompt / Context\n    └── Decision and generation: LLM\n\nInfrastructure: Model SDK / Agent framework / MCP\n</code></pre>\n<p>The same map can be compressed into a table:</p>\n<div class=\"table-scroll\"><table>\n<thead>\n<tr>\n<th>Concept</th>\n<th>Responsibility in the application</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>LLM</td>\n<td>Interpret input, make decisions, and generate output</td>\n</tr>\n<tr>\n<td>Prompt</td>\n<td>Describe the current task, rules, and output format</td>\n</tr>\n<tr>\n<td>Context</td>\n<td>Supply everything visible to the model on the current call</td>\n</tr>\n<tr>\n<td>RAG</td>\n<td>Retrieve external material and add it to generation</td>\n</tr>\n<tr>\n<td>Embedding / Vector DB</td>\n<td>Represent and search material by semantic similarity, as one RAG implementation path</td>\n</tr>\n<tr>\n<td>Fine-tuning</td>\n<td>Change model parameters and stable behavior through training</td>\n</tr>\n<tr>\n<td>Tool / Function Calling</td>\n<td>Let the model express an intent to use an external capability</td>\n</tr>\n<tr>\n<td>Workflow</td>\n<td>Arrange steps along a path defined in code</td>\n</tr>\n<tr>\n<td>Loop</td>\n<td>Repeat execution while a condition holds</td>\n</tr>\n<tr>\n<td>ReAct</td>\n<td>Let the model choose the next Action from Observations</td>\n</tr>\n<tr>\n<td>Memory / State</td>\n<td>Store information and task progress outside model calls</td>\n</tr>\n<tr>\n<td>Agent</td>\n<td>Combine the model, tools, state, loop, and boundaries into a running system</td>\n</tr>\n<tr>\n<td>MCP</td>\n<td>Connect a host to external capabilities through a standard protocol</td>\n</tr>\n<tr>\n<td>Agent framework</td>\n<td>Supply shared model integration, loop, and tool-execution infrastructure</td>\n</tr>\n</tbody>\n</table>\n</div><h2>Questions worth asking about an agent</h2>\n<p>When a feature calls itself an agent, the name is less useful than a few implementation questions:</p>\n<ol>\n<li>How many model calls can one task make, and what context does each call receive?</li>\n<li>Does code fix the retrieval sequence, or can the model choose another query from the result?</li>\n<li>Who executes tools, and where are arguments and permissions validated?</li>\n<li>Where is intermediate state stored, and can the task recover after failure?</li>\n<li>What stops the loop, and are there limits on steps, time, and cost?</li>\n<li>Which actions run automatically, and which ones require human approval?</li>\n</ol>\n<p>If a system cannot answer these questions, a demo may still work. Keeping it running in production will be harder. Agent engineering becomes difficult after the first model call, when context grows, tools fail, state must survive, and permissions begin to have consequences.</p>\n<h2>A real implementation</h2>\n<p>Concepts are more useful when they eventually meet code.</p>\n<p>I built a whole-book review feature for Novevia, an AI fiction-writing application. The first version sent the setting, chapters, facts, and memories to the model in one request. Two model calls took more than eight minutes. I later replaced that request with a manifest and let the model retrieve chapters, story routes, and facts as needed.</p>\n<p>That implementation contains RAG, tool use, and a bounded ReAct-style loop. Its task system also supplies state, stop conditions, read-only permissions, plan validation, human approval, recovery, and rollback. Calling the whole feature a constrained agent is reasonable.</p>\n<p><a href=\"/en/posts/2026/i-stopped-sending-the-whole-book-to-the-model/\">I Stopped Sending the Whole Book to the Model</a> records the implementation and its tradeoffs. It is one business-specific design, not the definition of these concepts. The concepts are a map for reading the code, not a set of prestigious labels to attach to it.</p>\n<p>The next time a feature calls itself an agent, look for three things: who chooses the next step, where state lives, and who makes it stop. Those answers say more than the name on the product page.</p>\n","date_published":"2026-08-08T00:00:00.000Z","tags":["AI","LLM","Agent","RAG","ReAct","Context Engineering"],"language":"en"},{"id":"https://www.lihuanyu.com/posts/context-engineering-is-not-more-context/","url":"https://www.lihuanyu.com/posts/context-engineering-is-not-more-context/","title":"上下文工程，不是把更多东西塞给模型","summary":"上下文窗口只规定模型最多能看到多少内容。真正的上下文工程，要处理材料的选择、压缩、排序、隔离和淘汰，还要给长任务留下可以继续工作的状态。","content_html":"<p>一个大模型功能很少在第一天就被上下文拖垮。</p>\n<p>第一版通常很清爽：一段 system prompt，一个用户问题，一次模型调用。后来答错了一次，就补一份说明；不知道业务资料，就接上 RAG；需要查实时状态，再加几个 Tool；为了让它记得前文，又把历史消息和任务日志一起带上。</p>\n<p>每一项单独看都合理。几个月后再抓一次请求，里面可能已经有规则、文档、搜索结果、工具返回、对话历史、用户偏好、失败记录和一段旧摘要。模型像坐在一张堆满文件的办公桌前。资料都在手边，真正要用的那页反而找不到了。</p>\n<p>这时再说“模型上下文窗口很大，塞得下”，往往已经答错了问题。</p>\n<p><img src=\"/assets/posts/2026/context-engineering/01-context-assembly-pipeline.jpg\" alt=\"散落的资料经过选择、压缩和排序，组成一轮模型真正使用的工作上下文\"></p>\n<p><a href=\"/en/posts/context-engineering-is-not-about-more-context/\">English version: Context Engineering Is Not About Adding More Context</a></p>\n<p><a href=\"/posts/from-one-model-call-to-an-agent/\">《从一次模型调用到一个 Agent》</a>介绍过 Prompt、Context、RAG、Memory 和 Agent 的关系。这篇往下走一步，只讨论一个更具体的问题：一轮模型调用之前，应用到底该把什么放进去，又该把什么留在外面。</p>\n<h2>窗口是容量，不是质量</h2>\n<p>上下文窗口（Context Window）规定一次推理最多能容纳多少 token。它像房间的面积，决定能搬进多少东西，却不保证东西摆得好，也不保证模型会看对地方。</p>\n<p>长上下文至少会带来四类麻烦。</p>\n<p>第一类很直接：输入越长，传输、预填充和推理通常越慢，费用也更高。Agent 每走一步都把前面的消息重新提交，单轮多出来的一点材料，会在后面的调用里反复付费。</p>\n<p>第二类是干扰。十份检索结果里只有两份相关，其余八份不是免费的陪衬，它们会和正确证据争夺注意力。过期规则、重复摘要和互相矛盾的记录尤其麻烦，模型往往会挑一个读起来顺的说法，而不是系统最希望它信的那一个。</p>\n<p>第三类是位置。<a href=\"https://arxiv.org/abs/2307.03172\">Lost in the Middle</a>一类研究指出，模型对长输入中不同位置的信息利用并不均匀。关键证据确实在上下文里，不代表它能被同样可靠地使用。把重要约束埋在中间，再用一句“请严格遵守以上要求”收尾，只能算一种愿望。</p>\n<p>第四类更容易被忽略：窗口还要容纳输出。输入把预算吃满以后，模型没有足够空间写答案、代码或结构化结果。有些接口会直接报错，有些运行时会从历史消息里截断一段。截的是闲话还是关键规则，要看运气。</p>\n<p>所以，长窗口解决的是“放不下”，上下文工程处理的是“该放什么、怎样摆、何时换”。</p>\n<h2>一轮调用里，究竟装了什么</h2>\n<p>Context 不只是用户输入的那句话。对一个带工具的应用来说，一轮调用常常包含这些部分：</p>\n<ul>\n<li>系统指令与安全规则</li>\n<li>当前任务和用户刚刚补充的要求</li>\n<li>从文档、数据库或搜索引擎取回的证据</li>\n<li>Tool 的返回结果</li>\n<li>最近几轮对话与任务状态</li>\n<li>少量长期记忆</li>\n<li>输出格式示例或 schema</li>\n</ul>\n<p>最后还要预留输出空间。</p>\n<p><img src=\"/assets/posts/2026/context-engineering/02-context-budget.jpg\" alt=\"上下文窗口被分成规则、当前任务、证据、工具结果、近期状态与输出预留\"></p>\n<p>这些材料的地位并不相同。系统规则通常稳定，当前任务必须完整，检索证据可以按相关性筛选，工具日志往往只需保留结论，历史对话则会随着任务推进逐步失效。把它们拼成一个字符串当然能跑，只是出了问题很难知道该删谁。</p>\n<p>更实用的做法是先分区，再给每个分区定预算和淘汰规则。预算不必一开始就精确到 token，可以先按字符数或消息条数做硬边界；但“所有材料共享一个总上限，超了就从最早消息开始砍”通常不够。最早那条消息里，可能正好放着任务目标。</p>\n<p>我更习惯先为输出留位置，再安排输入。规则和当前任务占固定区，证据与工具结果竞争弹性区，历史状态只保留继续工作需要的部分。这个顺序有点保守，却比模型写到 JSON 一半突然收工好得多。</p>\n<h2>上下文工程在做五件事</h2>\n<p>上下文工程不是一个新的 Prompt 技巧合集。Prompt 主要解决指令怎样写清楚；上下文工程面对的是材料从哪里来、怎样进入一轮调用，以及任务变长以后如何退出。</p>\n<p>落到代码里，大致是五个动作。</p>\n<p><strong>选择。</strong> 先根据当前问题找候选材料，再判断哪些值得进入工作集。用户问订单物流，不需要顺便附上退款制度全文；排查数据库超时，也不必把所有前端日志一起搬来。</p>\n<p><strong>压缩。</strong> 原始材料过长时，保留能支持判断的事实、来源和时间。压缩不是把十段文字改写成一段漂亮的废话。错误码、金额、版本号、文件路径这类信息一旦丢失，后面很难凭摘要找回来。</p>\n<p><strong>排序。</strong> 稳定规则放在稳定位置，当前任务明确标出，关键证据靠近需要使用它的地方。多份材料有时间或权威性差异时，顺序和标签都要说清楚，不能让模型自己猜哪一份更新。</p>\n<p><strong>隔离。</strong> 不同子任务使用不同工作集。查资料的模型不一定需要写作风格说明，负责生成最终答复的模型也不必看到工具调试日志。隔离能减少干扰，也能缩小敏感数据暴露的范围。</p>\n<p><strong>删除。</strong> 过期计划、重复 observation、已经被新结论替代的摘要，应当主动退出上下文。删除不是异常兜底，而是长任务的正常动作。只进不出的 Memory，最后只是另一种日志仓库。</p>\n<p>这五件事没有哪个听起来像魔法。真正麻烦的是，它们不能只在 prompt 模板里完成，还需要检索、状态存储、预算计算和运行时一起配合。</p>\n<h2>RAG、Memory 和 Tool Result 各有去处</h2>\n<p>几种常见材料容易混在一起，处理方式却不该一样。</p>\n<div class=\"table-scroll\"><table>\n<thead>\n<tr>\n<th>来源</th>\n<th>适合放什么</th>\n<th>进入 Context 的方式</th>\n<th>常见淘汰条件</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Prompt / Instructions</td>\n<td>目标、规则、输出格式</td>\n<td>每轮稳定保留，必要时分层</td>\n<td>任务结束或规则更新</td>\n</tr>\n<tr>\n<td>RAG</td>\n<td>与当前问题相关的外部事实</td>\n<td>按查询取回、去重并标注来源</td>\n<td>问题变化、证据过期或相关性下降</td>\n</tr>\n<tr>\n<td>Memory</td>\n<td>用户偏好、长期事实、历史结论</td>\n<td>先检索，再选少量加入</td>\n<td>被更新、失去可信度或不再相关</td>\n</tr>\n<tr>\n<td>Tool Result</td>\n<td>当前步骤刚取得的观察</td>\n<td>保留原始关键字段，旧结果逐步压缩</td>\n<td>已提取结论、被新结果替代</td>\n</tr>\n<tr>\n<td>Runtime State</td>\n<td>步数、阶段、失败次数、待办项</td>\n<td>用紧凑结构表达</td>\n<td>任务完成或进入新阶段</td>\n</tr>\n</tbody>\n</table>\n</div><p>RAG 的关键不只是“搜到几段”，而是搜什么、取几条、怎样去重，最后有没有把来源和时效带回来。Memory 也不是把所有历史聊天永久保存；保存只是入库，重新拿回当前任务才算使用。</p>\n<p>Tool Result 往往最容易失控。日志接口一次返回几百行，模型读完又调用第二个工具，运行时把两份原文连同解释全部追加到消息里。三四轮以后，真正的任务目标只占上下文的一小角。比较稳妥的做法是保留最近 observation 的原文，同时把较早结果压成带证据指针的阶段摘要。</p>\n<h2>长任务从第二次调用开始变难</h2>\n<p>一次调用的上下文可以手工拼出来。到了第二次，问题才真正出现：上一轮输出留多少，Tool Result 是否原样带回，失败尝试要不要保留，新证据和旧结论冲突时怎样处理。</p>\n<p>假设每轮都新增一段长度相近的 observation，而且下一轮把全部历史重新发送。第 10 轮付费的不是第 10 段，而是前 10 段的总和。整个任务的累计输入会近似按平方增长。循环没有边界时，窗口、延迟和账单会一起变胖。</p>\n<p>因此，长任务需要的不是一条不断加长的聊天记录，而是一套上下文生命周期。</p>\n<p><img src=\"/assets/posts/2026/context-engineering/03-long-task-context-lifecycle.jpg\" alt=\"长任务在检索、行动、观察、压缩、检查点和继续执行之间循环\"></p>\n<p>最近一两步可以保留细节，较早步骤压成 checkpoint。Checkpoint 至少要回答：目标是什么，已经确认了哪些事实，做过哪些动作，哪些尝试失败了，下一步还缺什么。它应当能让任务从这里恢复，而不是只写一句“前面进展顺利”。</p>\n<p>子任务完成后，它的过程材料可以退出主上下文，只把结论、证据和未决问题交回来。遇到新的阶段，再重新检索需要的资料。这样做看似多了几次装配，实际上让每轮调用有了清楚的工作台。</p>\n<h2>一轮 Context 可以怎样装配</h2>\n<p>实现不必从复杂框架开始。先让“收集材料”和“调用模型”变成两个明确步骤，很多问题已经会暴露出来。</p>\n<pre><code class=\"language-ts\">const candidates = await collectCandidates(task, state)\nconst evidence = selectAndRank(candidates, task)\nconst recent = keepRecentObservations(state.events, 2)\nconst checkpoint = compactOlderEvents(state.events)\n\nconst context = assemble({\n  instructions: stableInstructions,\n  task: normalizeTask(task),\n  evidence: fitToBudget(evidence, budgets.evidence),\n  recent,\n  checkpoint,\n  outputSchema,\n  reserveForOutput: budgets.output,\n})\n\nconst result = await model.generate(context)\nawait persistResultAndState(result, state)\n</code></pre>\n<p>这里最值得留下的不是函数名，而是边界：候选材料不等于最终 Context；近期 observation 和历史状态使用不同保留策略；输出预算在调用前就参与装配；模型返回以后，结果和状态要保存到 Context 之外。</p>\n<p>再往前走，可以给每段材料附上 <code>source</code>、<code>timestamp</code>、<code>priority</code>、<code>ttl</code> 和 <code>sensitivity</code>。它来自哪里、什么时候产生、多久过期、能不能发送给当前模型，都不该只藏在正文里。</p>\n<h2>怎么知道上下文改对了</h2>\n<p>“Token 变少了”是一个指标，不是最后答案。删得太狠，模型会更快地答错。</p>\n<p>我会同时看三组数据。</p>\n<p>第一组是成本与延迟：每轮输入 token、整个任务累计 token、首 token 时间、总耗时，以及 p50 和 p95。平均值很容易把少数特别肥的任务藏起来。</p>\n<p>第二组是任务质量：固定一批真实问题，检查答案是否引用了正确证据，结构化输出能否通过校验，任务完成率和人工采纳率有没有变化。对于 Tool，还要看模型取了哪些资料，而不只是最终答案读起来顺不顺。</p>\n<p>第三组是运行过程：检索了多少候选、最终选入多少、截断发生在哪个分区、摘要被复用了几轮、任务为什么停止。没有这些记录，Context 出错时只剩一句“模型不稳定”，这句话通常什么也修不了。</p>\n<p>评估集不必很大，先收集二三十个真正出过问题的任务就有价值。重要的是固定输入、期望证据和通过条件。否则每换一个模型或策略，都只能靠聊天窗口里的几次试用下结论。</p>\n<h2>几个常见误区</h2>\n<p>最常见的是把“模型可能用得上”当成“每轮都要带上”。可能有用的材料应该留在可检索的存储里，只有当前问题需要时才进入 Context。</p>\n<p>第二个误区是过早总结。为了省 token，系统先把原文压成没有出处的概括，后面再也找不回数字和细节。摘要应该带来源指针；高风险事实最好保留原文片段，必要时允许模型重新读取。</p>\n<p>第三个误区是给每条材料同样的信任。用户刚提供的信息、数据库实时查询、三个月前的聊天摘要和模型自己上轮的猜测，不该排成同一级。来源、时效和置信度必须显式存在。</p>\n<p>还有一种做法是只在超限时报错，然后粗暴截掉最早消息。它能让接口继续返回 200，却可能悄悄删掉目标和约束。真正有用的降级策略应该按分区执行：先去重，再删低相关证据，再压缩旧 observation，最后才考虑缩减稳定指令。</p>\n<h2>我在一个长任务里踩过的坑</h2>\n<p>我做的小说应用 Novevia 有一个“整理全书”功能。早期版本把作品设定、章节卡片、事实和记忆尽量压缩后，一次性发给模型。书越长，第一个请求越大，两次模型调用曾跑到八分多钟。</p>\n<p>后来我没有继续研究怎样写一个更厉害的大 Prompt，而是改变了 Context 的装配方式。第一轮只给作品 manifest，模型缺什么再查章节、路线或事实；每次请求设独立预算，过长时保留规则、manifest 和近期 Tool Result；任务状态留在数据库，模型只看到继续判断所需的部分。</p>\n<p>它也因此形成了一个有限的 ReAct-style Loop，但 ReAct 不是解决超时的咒语。真正起作用的是工作集变小了，检索有了时机，历史开始淘汰，循环也知道何时停。</p>\n<p><a href=\"/posts/ai-book-review-beyond-large-prompt/\">《AI 整理全书为什么不能只靠一个大 Prompt》</a>记录了这段实现，包括只读工具、步数限制、校验、恢复和回滚。那是一个具体项目；放到别的大模型应用里，资料名称会变，装配问题不会消失。</p>\n<p>上下文不能太穷，缺了证据，模型只能猜。也不能把仓库搬进窗口，再期待模型自己整理货架。</p>\n<p>比较实际的标准只有一个：每轮调用都让模型看到完成当前一步所需的材料，同时没有多到妨碍它工作。做到这一点，才算真正开始做上下文工程。</p>\n","date_published":"2026-08-08T00:00:00.000Z","tags":["AI","LLM","上下文工程","RAG","Agent"],"language":"zh"},{"id":"https://www.lihuanyu.com/posts/from-one-model-call-to-an-agent/","url":"https://www.lihuanyu.com/posts/from-one-model-call-to-an-agent/","title":"从一次模型调用到一个 Agent","summary":"LLM、Prompt、Context、RAG、Tool、Workflow、Loop、ReAct、Memory 和 Agent 经常出现在同一段介绍里，却不处于同一层：模型负责生成，上下文与 RAG 补充材料，Tool 连接外部能力，Loop 和 ReAct 组织决策，Agent 把它们变成一套运行系统。","content_html":"<p>接入一个模型 API，应用可以对话；再接一个搜索接口，介绍里开始出现 RAG；让模型调用几个工具，Agent 也来了；外面再套一层 <code>while</code>，有时连 ReAct 都一起写上。</p>\n<p>名词涨得比代码快。</p>\n<p>麻烦在于，这些词并不处于同一个层面。LLM 是模型，Prompt 和 Context 是输入，RAG 和 Tool 给模型补充外部能力，Workflow、Loop 和 ReAct 负责组织执行过程，Agent 才是把它们装起来运行的系统。MCP 和 Pi 这类协议或框架，又在更外面一层。</p>\n<p>判断一个大模型应用，比较有用的办法不是先看它叫什么，而是看四件事：模型每轮能看到什么，下一步由谁决定，外部操作由谁执行，任务最后怎样停下。</p>\n<p>从这四个问题出发，概念之间的关系会清楚许多。</p>\n<p><img src=\"/assets/posts/2026/llm-agent-concepts/01-llm-agent-concept-map.jpg\" alt=\"LLM、上下文、外部能力、编排、状态与 Agent 边界组成的分层概念图\"></p>\n<p><a href=\"/en/posts/2026/from-one-model-call-to-an-agent/\">English version: From One Model Call to an Agent</a></p>\n<h2>LLM 只是模型，不是整个应用</h2>\n<p>大语言模型（Large Language Model，LLM）是大模型应用里的生成与决策部件。</p>\n<p>从训练原理看，它根据已有 token 预测后续 token；从应用开发的角度看，它接收一组消息，再生成文本或结构化输出。应用调用模型完成生成的过程，通常叫推理（Inference）。</p>\n<p>最小的调用可以只有一句要求：</p>\n<pre><code class=\"language-text\">请用三句话总结下面这段内容。\n</code></pre>\n<p>模型收到要求和待总结的文字，返回结果，调用结束。这就是一次普通的 LLM 调用。模型不需要工具，不需要循环，也不需要 Agent。</p>\n<p>模型本身和完整应用不是一回事。页面、用户身份、数据库、权限、历史消息、重试和计费都由外围系统负责。即使聊天界面看起来记得前文，也是应用或模型服务保存了对话状态，并在后续推理中提供相关内容。</p>\n<p>从模型本身看，当前输出取决于当前可用的上下文。模型参数不会因为昨天回答过某个用户，今天就自动多出一段针对这个用户的记忆。</p>\n<h2>Prompt 决定任务，Context 决定模型看见什么</h2>\n<p>Prompt 和 Context 经常被混在一起，但两者范围不同。</p>\n<p>Prompt 主要描述任务：模型扮演什么角色，要完成什么目标，遵守哪些规则，按什么格式回答。System Prompt、用户问题和输出格式约束都属于这一层。</p>\n<p>不同 API 对 Prompt 一词的用法并不统一，有时它也泛指整包模型输入。为了把职责说清楚，这里采用较窄的定义：Prompt 表达任务，Context 包含模型本轮能看到的一切。</p>\n<p>Context 的范围更大。只要某段内容进入本轮模型输入，它就是上下文的一部分，例如：</p>\n<ul>\n<li>系统指令和用户问题</li>\n<li>需要处理的原始材料</li>\n<li>对话历史</li>\n<li>检索回来的文档</li>\n<li>可用工具的说明</li>\n<li>前几轮工具返回的结果</li>\n</ul>\n<p>上下文窗口（Context Window）规定一次调用最多能容纳多少 token。Token 是模型处理文本的基本单位，可能是一个汉字、半个英文单词或一个标点，不能直接和字符数画等号。</p>\n<p>Prompt 工程更关心指令怎样表达。上下文工程（Context Engineering）还要决定材料从哪里来、什么时候加入、保留多久，以及超过预算时先删什么。</p>\n<p>可以把 LLM 想成坐在桌前处理材料的人。Prompt 是任务单，Context 是桌上全部资料，Context Window 是桌面大小。桌子更大当然能放下更多东西，但把资料室里的纸全搬上来，不会自动得到更好的判断。</p>\n<h2>RAG 给生成过程补充外部资料</h2>\n<p>模型训练完成以后，不会自动知道企业私有文档、刚发生的新闻或数据库里的最新记录。<a href=\"https://arxiv.org/abs/2005.11401\">检索增强生成（Retrieval-Augmented Generation，RAG）</a>用外部检索补上这部分信息。</p>\n<p>一个常见问题是：</p>\n<blockquote>\n<p>公司的差旅报销上限是多少？</p>\n</blockquote>\n<p>模型可以凭训练数据猜一个答案，也可以先从公司制度里查到相关条款，再依据条款回答。后者的基本流程是：</p>\n<pre><code class=\"language-text\">问题 -&gt; 检索相关资料 -&gt; 加入上下文 -&gt; 调用模型 -&gt; 生成答案\n</code></pre>\n<p>RAG 改变的是模型本轮收到的资料，不是模型参数。制度文件更新以后，只需更新资料库，不必重新训练模型。</p>\n<p>这也是 RAG 和 Fine-tuning 的主要区别。RAG 在推理时补充事实，适合持续变化的知识；Fine-tuning 通过训练调整模型参数，更适合改变稳定的行为、风格或特定任务能力。两者可以一起用，但解决的不是同一个问题。</p>\n<p>RAG 也不等于向量数据库。Embedding 会把文本转换成便于比较语义距离的向量，向量数据库负责保存和检索这些向量。它们是语义检索的常用实现，却不是 RAG 的定义。</p>\n<p>全文搜索、关键词匹配、SQL 查询，甚至按文件路径读取文档，也能为生成过程补充外部资料。只要系统先取回相关信息，再把信息用于生成，就具备 RAG 的基本结构。</p>\n<p>固定 RAG 的检索步骤通常由代码安排。服务端收到问题，生成查询，取回前几条结果，再调用一次模型。整个过程可以只有一个 Workflow，不需要 Agent，也不需要 ReAct。</p>\n<h2>Tool 让模型请求外部能力</h2>\n<p>RAG 主要解决“回答前缺少资料”，Tool 的范围更大。</p>\n<p>Tool 可以读取信息，也可以执行动作。查询订单、搜索日志、读取文件属于读操作；发送邮件、创建工单、修改日程和执行部署属于写操作。模型本身通常不直接拥有这些权限，它只能表达调用意图。</p>\n<p>常见的 Function Calling 会把工具名称、用途和参数结构告诉模型。用户问“订单什么时候发货”，模型可以返回：</p>\n<pre><code class=\"language-json\">{\n  &quot;tool&quot;: &quot;get_order&quot;,\n  &quot;args&quot;: { &quot;orderId&quot;: &quot;1234567890123&quot; }\n}\n</code></pre>\n<p>这段 JSON 不会自己查询数据库。应用需要验证订单号和当前用户权限，执行 <code>get_order</code>，再把结果交给模型组织回答。</p>\n<p>模型负责提出“想调用什么”，运行时负责决定“能不能调用”和“怎样执行”。这个分工很重要。否则只要模型生成了一段看起来像命令的文本，系统就得照着做，迟早会出事。</p>\n<p>Tool 和 RAG 会有交集。把“搜索知识库”做成 Tool，模型调用它以后，结果进入生成过程，这同时是 Tool Use 和 RAG。发送邮件也是 Tool Use，却通常不会被叫作 RAG，因为重点是改变外部系统，不是补充回答资料。</p>\n<p>调用一次 Tool 也不必然构成 Agent。固定地查询一次订单、生成一次回复，仍然可以是一段普通 Workflow。</p>\n<h2>Workflow、Loop 和 ReAct 回答三个问题</h2>\n<p>Workflow、Loop 和 ReAct 经常被放在一起比较，其实它们回答的是三个不同问题。</p>\n<p><img src=\"/assets/posts/2026/llm-agent-concepts/02-workflow-loop-react.jpg\" alt=\"Workflow 预设路径、Loop 按条件重复、ReAct 根据观察选择下一步\"></p>\n<p>Workflow 关心步骤由谁安排。比如一套故障报告流程由代码规定：先查服务状态，再读取最近错误日志，然后汇总部署记录，最后交给模型生成报告。路径已经写好，模型只处理其中的内容。</p>\n<p>Loop 关心过程是否重复。重试三次是 Loop，轮询任务状态是 Loop，模型连续修改代码直到测试通过也是 Loop。循环只是一种控制结构，本身没有“智能”。</p>\n<p><a href=\"https://arxiv.org/abs/2210.03629\">ReAct</a>关心模型在循环里怎样选择下一步。这个名字来自 Reasoning and Acting：模型根据当前信息作出判断，选择一个 Action，看到 Observation，再继续判断。</p>\n<p>同样是排查故障，ReAct-style Loop 可能走出这样的路径：</p>\n<pre><code class=\"language-text\">观察告警\n-&gt; 决定查询错误日志\n-&gt; 发现数据库连接超时\n-&gt; 决定读取连接池配置\n-&gt; 发现配置刚被修改\n-&gt; 决定查询部署记录\n-&gt; 给出诊断结论\n</code></pre>\n<p>外层代码负责继续或终止循环，模型根据每轮 Observation 选择下一个 Action。查询顺序不是服务端提前写死的，而是由已经取得的证据决定。</p>\n<p>Loop 不等于 ReAct。固定 Workflow 可以放在 Loop 里反复执行，重试循环也不需要模型参与。ReAct 只是模型驱动循环的一种决策方式，也不是构建 Agent 的唯一办法。</p>\n<p>实现 ReAct 也不要求保存完整的长思维过程。运行时记录简短的行动理由、工具参数和观察结果，通常已经足够还原任务路径。真正需要审计的是模型做了什么、系统返回了什么，以及哪条规则允许它继续。</p>\n<h2>Agent 把模型、工具和运行时装在一起</h2>\n<p>Agent 没有一条所有人都接受的分界线。有些产品只要调用过 Tool，就愿意在名称里加上 Agent。工程实现仍然需要一个能落到代码上的定义。</p>\n<p>可以把 Agent 理解为一套围绕目标持续决策和行动的运行系统。一个可用的 Agent 通常能找到这些部件：</p>\n<ul>\n<li><strong>Model</strong>：理解当前状态并决定下一步</li>\n<li><strong>Instructions 与 Context</strong>：提供目标、规则和本轮材料</li>\n<li><strong>Tools</strong>：读取信息或影响外部系统</li>\n<li><strong>State</strong>：保存任务进度和已经发生的事件</li>\n<li><strong>Control Loop</strong>：组织模型、Action 和 Observation 之间的往返</li>\n<li><strong>Stop Conditions</strong>：定义完成、失败、超时和人工接管</li>\n<li><strong>Boundaries</strong>：限制权限、预算、步数和危险操作</li>\n</ul>\n<p>写成一个不太严谨但容易记住的式子，就是：</p>\n<pre><code class=\"language-text\">Agent = Model + Context + Tools + State + Loop + Boundaries\n</code></pre>\n<p><img src=\"/assets/posts/2026/llm-agent-concepts/03-bounded-agent-system.jpg\" alt=\"Agent 的模型、工具、状态、运行时、停止条件和权限边界\"></p>\n<p>以线上故障诊断为例，Agent 的目标是找出告警原因。它可以读取监控、日志、配置和部署记录，在多轮查询中保留线索，最后提交诊断方案。代码还要限制最大步数、总时间和可用工具，写操作则停下来等待人工确认。</p>\n<p>“越自主越好”不是 Agent 的目标。生产系统往往需要更多边界：只读工具、调用预算、超时、审批、校验、幂等和回滚。这些工作不新鲜，却决定 Agent 出错时会留下一个失败任务，还是留下一场事故。</p>\n<p>Workflow 和 Agent 也不是非黑即白。代码决定得越多，系统越接近固定 Workflow；模型能根据状态选择的路径越多，系统越 agentic。可靠的实现通常会把两者组合起来：模型负责探索，固定流程负责验证、审批和执行。</p>\n<h2>Memory 不等于更大的上下文窗口</h2>\n<p>Agent 跨越多轮甚至多个任务工作时，需要把一部分信息保存在模型调用之外。</p>\n<p>Context 是模型本轮能看到的内容，Memory 则是应用长期保存、以后可能再次取用的信息。Memory 可以放在数据库、文件或专门的存储系统里，但保存完成不代表模型已经知道。应用仍要在合适的时候读取它，再放回当前 Context。</p>\n<p>应用里的“记忆”至少有三类：</p>\n<ul>\n<li><strong>Conversation state</strong>：当前对话说过什么</li>\n<li><strong>Domain memory</strong>：用户偏好、业务事实、文档摘要等长期信息</li>\n<li><strong>Runtime state</strong>：任务走到哪一步、调用过哪些工具、失败过几次</li>\n</ul>\n<p>把所有历史消息一直带着，不等于设计好了 Memory。那只是让 Context 越来越长。真正的 Memory 还要决定写入什么、什么时候读取、信息冲突时信谁，以及过期内容怎么处理。</p>\n<h2>MCP、Pi 和 Agent framework 属于基础设施</h2>\n<p>模型、输入、工具和编排方式分清以后，协议与框架的位置也容易判断。</p>\n<p><a href=\"https://modelcontextprotocol.io/docs/getting-started/intro\">模型上下文协议（Model Context Protocol，MCP）</a>提供一种标准方式，让宿主连接外部的 Tool、Resource 和 Prompt。它解决的是能力怎样暴露和连接，不会替 Agent 决定任务目标、循环策略和权限边界。</p>\n<p>Pi、LangChain、Mastra 一类项目则会处理一部分通用工程工作，例如统一模型接口、组织消息、声明工具、运行循环、记录 trace，或者接入不同 provider。每个项目覆盖的层次不同，不能只看“Agent Framework”这个名字。</p>\n<p>以我在项目中评估过的 Pi 为例，<code>@earendil-works/pi-ai</code> 更接近模型接入层，负责 provider、消息、流式输出、tool call 和 usage 的统一。采用相应的 Agent runtime，还能少写一部分循环与工具执行代码。领域规则仍然留在应用里：哪些数据可以读，哪些动作必须确认，什么结果才算完成。</p>\n<p>框架像运行底盘，MCP 像连接标准。它们能减少通用代码，但不会替产品完成领域判断。</p>\n<h2>把这些概念放在同一张图里</h2>\n<p>按大模型应用的运行层次，可以这样理解这些概念：</p>\n<pre><code class=\"language-text\">应用目标与边界\n└── Agent\n    ├── 编排：Workflow / Loop / ReAct\n    ├── 外部能力：RAG / Tool\n    ├── 状态：Memory / Runtime State\n    ├── 本轮输入：Prompt / Context\n    └── 决策与生成：LLM\n\n基础设施：Model SDK / Agent framework / MCP\n</code></pre>\n<p>它们各自负责的事情也可以压缩成一张表：</p>\n<div class=\"table-scroll\"><table>\n<thead>\n<tr>\n<th>概念</th>\n<th>在应用里负责什么</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>LLM</td>\n<td>理解输入、作出判断、生成输出</td>\n</tr>\n<tr>\n<td>Prompt</td>\n<td>描述本轮任务、规则和输出格式</td>\n</tr>\n<tr>\n<td>Context</td>\n<td>提供模型本轮能看到的全部材料</td>\n</tr>\n<tr>\n<td>RAG</td>\n<td>从外部资料中取回相关内容，加入生成过程</td>\n</tr>\n<tr>\n<td>Embedding / Vector DB</td>\n<td>按语义表示和检索资料，是 RAG 的一种实现路径</td>\n</tr>\n<tr>\n<td>Fine-tuning</td>\n<td>通过训练调整模型参数和稳定行为</td>\n</tr>\n<tr>\n<td>Tool / Function Calling</td>\n<td>让模型表达调用外部能力的意图</td>\n</tr>\n<tr>\n<td>Workflow</td>\n<td>按代码预先定义的路径组织步骤</td>\n</tr>\n<tr>\n<td>Loop</td>\n<td>让执行过程按条件重复</td>\n</tr>\n<tr>\n<td>ReAct</td>\n<td>让模型根据观察结果选择下一次行动</td>\n</tr>\n<tr>\n<td>Memory / State</td>\n<td>在模型调用之外保存信息和任务进度</td>\n</tr>\n<tr>\n<td>Agent</td>\n<td>把模型、工具、状态、循环和边界组织成运行系统</td>\n</tr>\n<tr>\n<td>MCP</td>\n<td>用标准协议连接宿主与外部能力</td>\n</tr>\n<tr>\n<td>Agent framework</td>\n<td>提供模型接入、循环、工具执行等通用基础设施</td>\n</tr>\n</tbody>\n</table>\n</div><h2>判断一个 Agent 时，可以先问什么</h2>\n<p>面对一个自称 Agent 的功能，不必急着争论名字。先看实现里有没有几个明确答案：</p>\n<ol>\n<li>模型一共会被调用几次，每轮能看到哪些 Context？</li>\n<li>检索内容由代码固定，还是模型根据结果继续选择？</li>\n<li>Tool 由谁执行，参数和权限在哪里校验？</li>\n<li>中间状态存在哪里，失败后能不能恢复？</li>\n<li>Loop 在什么条件下结束，有没有步数、时间和费用上限？</li>\n<li>哪些动作可以自动完成，哪些动作必须让人确认？</li>\n</ol>\n<p>如果这些问题没有答案，系统即使能演示，也很难长期运行。Agent 的难处很少在第一次调用模型，而在第二轮以后：上下文开始膨胀，工具可能失败，状态需要保存，权限也开始真正产生后果。</p>\n<h2>一个真实的落地例子</h2>\n<p>概念最终还是要回到代码。</p>\n<p>我在小说应用 Novevia 里做过一个全书整理功能。最早的版本把作品设定、章节、事实和记忆一次性发给模型，两次调用跑了八分多钟。后来它改成先给 manifest，再由模型按需查询章节、路线和事实库。</p>\n<p>这个实现同时包含 RAG、Tool Use 和一个有限的 ReAct-style Loop。任务系统还负责状态、停止条件、只读权限、方案校验、人工确认、恢复和回滚，因此整个功能也可以称为一个受约束的 Agent。</p>\n<p><a href=\"/posts/ai-book-review-beyond-large-prompt/\">《AI 整理全书为什么不能只靠一个大 Prompt》</a>记录了那次改造的完整过程。它只是一种业务实现，不是这些概念的定义；反过来，这些概念也只是解释代码的地图，不是给功能贴金的称号。</p>\n<p>下次再看到一个功能自称 Agent，可以先找三样东西：下一步由谁决定，状态存在哪里，谁负责让它停手。三个问题，比产品页上的名字有用。</p>\n","date_published":"2026-08-08T00:00:00.000Z","tags":["AI","LLM","Agent","RAG","ReAct","上下文工程"],"language":"zh"},{"id":"https://www.lihuanyu.com/en/posts/2026/run-redis-with-docker-compose-on-ubuntu/","url":"https://www.lihuanyu.com/en/posts/2026/run-redis-with-docker-compose-on-ubuntu/","title":"How to Run Redis with Docker Compose on Ubuntu","summary":"Deploy Redis 8 on Ubuntu with Docker Compose using localhost-only access, authentication, AOF persistence, memory limits, backups, and safe upgrades.","content_html":"<p>I first ran Redis in Docker on a small Ubuntu server to synchronize configuration between services. Starting the container took one command. Deciding what should happen to the port, password, memory, logs, and data took much longer.</p>\n<p>This setup is for a small Redis instance used as a cache or lightweight application dependency on a single Linux host. It keeps Redis reachable only from that host, persists data with AOF, limits memory, rotates container logs, and includes a backup and upgrade path.</p>\n<p>If the main question is whether Docker itself is too expensive for a small server, read <a href=\"/en/posts/2025/rethinking-docker-development-linux-redis/\">Docker Performance on Linux vs Docker Desktop</a>. Docker Engine on Linux has a very different resource profile from Docker Desktop on macOS or Windows.</p>\n<h2>Decide how Redis will be reached</h2>\n<p>The safest network configuration depends on where the application runs:</p>\n<div class=\"table-scroll\"><table>\n<thead>\n<tr>\n<th>Application location</th>\n<th>Redis connection</th>\n<th>Compose port setting</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>On the same Ubuntu host</td>\n<td><code>127.0.0.1:6379</code></td>\n<td>Bind <code>127.0.0.1:6379:6379</code></td>\n</tr>\n<tr>\n<td>In the same Compose project</td>\n<td><code>redis:6379</code></td>\n<td>Do not publish a port</td>\n</tr>\n<tr>\n<td>On another server</td>\n<td>Private network or SSH tunnel</td>\n<td>Do not publish Redis to the public internet</td>\n</tr>\n</tbody>\n</table>\n</div><p>The example below assumes the application runs directly on the same Ubuntu host. Binding the port to <code>127.0.0.1</code> prevents external clients from reaching it through the server’s public interface.</p>\n<p>This matters even when <code>ufw</code> is enabled. Docker’s firewall documentation warns that published container ports can bypass some host firewall expectations. A loopback or private-network binding is a better first boundary than a public bind plus a firewall rule.</p>\n<h2>Prerequisites</h2>\n<p>Use a currently supported Ubuntu release and install Docker Engine from Docker’s official apt repository. The exact repository commands can change, so follow the <a href=\"https://docs.docker.com/engine/install/ubuntu/\">current Docker Engine installation guide</a> instead of copying an old package source indefinitely.</p>\n<p>Verify both Docker Engine and the Compose plugin:</p>\n<pre><code class=\"language-bash\">docker --version\ndocker compose version\n</code></pre>\n<p>Access to the Docker socket is effectively root-level access to the host. Only add trusted administrative users to the <code>docker</code> group.</p>\n<h2>Create the service directory and password</h2>\n<p>Create a dedicated directory:</p>\n<pre><code class=\"language-bash\">sudo install -d -m 0750 -o &quot;$USER&quot; -g &quot;$USER&quot; /opt/redis\ncd /opt/redis\nmkdir -p data\n</code></pre>\n<p>Generate a password in a local <code>.env</code> file:</p>\n<pre><code class=\"language-bash\">umask 077\nprintf 'REDIS_PASSWORD=%s\\n' &quot;$(openssl rand -base64 32)&quot; &gt; .env\nchmod 600 .env\n</code></pre>\n<p>Do not commit <code>.env</code>. This keeps the password out of the Compose file and shell history, but it is not a full secret manager. A user with root or Docker socket access can still inspect the running container. For a multi-user production platform, use the platform’s secret-management system and Redis ACLs.</p>\n<h2>Create <code>compose.yaml</code></h2>\n<p>Create this file under <code>/opt/redis</code>:</p>\n<pre><code class=\"language-yaml\">services:\n  redis:\n    image: redis:8-alpine\n    restart: unless-stopped\n    ports:\n      - &quot;127.0.0.1:6379:6379&quot;\n    command:\n      - redis-server\n      - --appendonly\n      - &quot;yes&quot;\n      - --appendfsync\n      - everysec\n      - --maxmemory\n      - &quot;64mb&quot;\n      - --maxmemory-policy\n      - allkeys-lru\n      - --requirepass\n      - &quot;${REDIS_PASSWORD:?set REDIS_PASSWORD in .env}&quot;\n    environment:\n      REDIS_PASSWORD: &quot;${REDIS_PASSWORD:?set REDIS_PASSWORD in .env}&quot;\n    volumes:\n      - ./data:/data\n    mem_limit: 128m\n    healthcheck:\n      test:\n        - CMD-SHELL\n        - 'REDISCLI_AUTH=&quot;$$REDIS_PASSWORD&quot; redis-cli ping | grep -q PONG'\n      interval: 10s\n      timeout: 3s\n      retries: 5\n    logging:\n      driver: json-file\n      options:\n        max-size: &quot;10m&quot;\n        max-file: &quot;3&quot;\n</code></pre>\n<p>Several choices here are deliberate:</p>\n<ul>\n<li><code>redis:8-alpine</code> pins the major Redis version while still receiving compatible patch releases.</li>\n<li><code>127.0.0.1</code> keeps the published port on the local host.</li>\n<li>AOF with <code>appendfsync everysec</code> provides a practical durability/performance balance for a small service.</li>\n<li><code>maxmemory 64mb</code> bounds Redis data, while <code>mem_limit: 128m</code> bounds the whole container.</li>\n<li><code>allkeys-lru</code> is suitable when Redis is a cache and old keys may be evicted.</li>\n<li>The health check authenticates without putting the password directly in the command.</li>\n<li>Log rotation prevents a noisy container from filling the server disk indefinitely.</li>\n</ul>\n<p>The Redis data limit and container limit are intentionally different. Redis needs memory for client connections, replication buffers, allocator overhead, persistence work, and other processes beyond stored keys. Do not set the container limit equal to <code>maxmemory</code>.</p>\n<p>If Redis stores jobs, configuration, or anything that must not be evicted, replace <code>allkeys-lru</code> with <code>noeviction</code> and size the instance for the expected dataset. Redis persistence improves recovery, but it does not turn an eviction cache into a primary database.</p>\n<h2>Validate and start Redis</h2>\n<p>Validate the Compose file without printing the interpolated configuration:</p>\n<pre><code class=\"language-bash\">docker compose config --quiet\n</code></pre>\n<p>Then start Redis:</p>\n<pre><code class=\"language-bash\">docker compose up -d\ndocker compose ps\n</code></pre>\n<p>The status should become <code>healthy</code>. Test Redis from inside the container without typing the password into the command:</p>\n<pre><code class=\"language-bash\">docker compose exec redis sh -lc \\\n  'REDISCLI_AUTH=&quot;$REDIS_PASSWORD&quot; redis-cli ping'\n</code></pre>\n<p>The expected response is:</p>\n<pre><code class=\"language-text\">PONG\n</code></pre>\n<p>Inspect memory and persistence state:</p>\n<pre><code class=\"language-bash\">docker compose exec redis sh -lc \\\n  'REDISCLI_AUTH=&quot;$REDIS_PASSWORD&quot; redis-cli INFO memory'\n\ndocker compose exec redis sh -lc \\\n  'REDISCLI_AUTH=&quot;$REDIS_PASSWORD&quot; redis-cli INFO persistence'\n\ndocker stats --no-stream\n</code></pre>\n<p>The first startup also creates AOF data under <code>/opt/redis/data</code>. Check that the directory is not empty before assuming persistence is working.</p>\n<h2>Connect an application on the same host</h2>\n<p>Load the password into the current shell only for the process that needs it:</p>\n<pre><code class=\"language-bash\">set -a\n. /opt/redis/.env\nset +a\nREDISCLI_AUTH=&quot;$REDIS_PASSWORD&quot; redis-cli -h 127.0.0.1 ping\nunset REDIS_PASSWORD\n</code></pre>\n<p>Application libraries usually accept a URL in this shape:</p>\n<pre><code class=\"language-text\">redis://:PASSWORD@127.0.0.1:6379/0\n</code></pre>\n<p>Keep that URL in the application’s secret configuration, not in source control or client-side code.</p>\n<p>If the application runs in the same Compose project, remove the <code>ports</code> block and connect to <code>redis:6379</code> over the Compose network. That removes the host port entirely.</p>\n<h2>Remote access without a public Redis port</h2>\n<p>For occasional administration from another machine, an SSH tunnel is safer than opening port 6379:</p>\n<pre><code class=\"language-bash\">ssh -L 6379:127.0.0.1:6379 user@example-server\n</code></pre>\n<p>While the tunnel is open, a local Redis client can connect to <code>127.0.0.1:6379</code>. For a permanent connection between servers, use a private network, restrict source addresses at the cloud network layer, enable TLS where required, and prefer Redis ACL users with only the permissions each application needs.</p>\n<p>Password authentication alone is not a reason to expose Redis directly to the internet.</p>\n<h2>Back up the data</h2>\n<p>AOF protects against normal restarts, but files on the same server are not a backup. Export an RDB snapshot and copy it outside the container:</p>\n<pre><code class=\"language-bash\">mkdir -p backups\n\ndocker compose exec redis sh -lc \\\n  'REDISCLI_AUTH=&quot;$REDIS_PASSWORD&quot; redis-cli --rdb /tmp/redis-backup.rdb'\n\ndocker compose cp \\\n  redis:/tmp/redis-backup.rdb \\\n  &quot;./backups/redis-$(date +%F).rdb&quot;\n\ndocker compose exec redis rm -f /tmp/redis-backup.rdb\n</code></pre>\n<p>Move the resulting backup to storage outside this server and test restoration before relying on it. A backup that has never been restored is only a hopeful file.</p>\n<h2>Upgrade Redis safely</h2>\n<p>Because the image pins Redis 8 rather than a patch release, <code>docker compose pull</code> can bring in a newer Redis 8 image. Before upgrading:</p>\n<ol>\n<li>Create and copy a backup.</li>\n<li>Read the Redis image and server release notes.</li>\n<li>Record the currently running image digest.</li>\n<li>Pull and recreate the container.</li>\n<li>Verify health, logs, persistence, and an application read/write path.</li>\n</ol>\n<p>The update commands are:</p>\n<pre><code class=\"language-bash\">docker compose pull redis\ndocker compose up -d\ndocker compose ps\ndocker compose logs --tail=100 redis\n</code></pre>\n<p>For a major-version upgrade, review compatibility and persistence format changes first. Do not change the image tag and the application client at the same time unless there is a tested rollback plan.</p>\n<h2>Routine checks</h2>\n<p>A small Redis instance does not need a large operations platform, but it should not be invisible. Check these periodically:</p>\n<pre><code class=\"language-bash\">docker compose ps\ndocker compose logs --tail=100 redis\ndocker stats --no-stream\ndf -h\ndu -sh /opt/redis/data\n</code></pre>\n<p>Also watch application-visible signals: connection errors, command latency, evicted keys, rejected writes under <code>noeviction</code>, and failed persistence operations.</p>\n<h2>Conclusion</h2>\n<p>Running Redis with Docker Compose is easy. Running it with clear failure boundaries takes a little more work.</p>\n<p>For a small Ubuntu host, the useful defaults are simple: keep Redis off the public interface, authenticate clients, persist data, separate Redis memory from the container limit, rotate logs, export backups, and plan upgrades. Docker makes these choices repeatable, but it does not choose them for you.</p>\n<h2>Further reading</h2>\n<ul>\n<li><a href=\"/en/posts/2025/rethinking-docker-development-linux-redis/\">Docker Performance on Linux vs Docker Desktop</a></li>\n<li><a href=\"https://docs.docker.com/engine/install/ubuntu/\">Docker Docs: Install Docker Engine on Ubuntu</a></li>\n<li><a href=\"https://docs.docker.com/engine/network/packet-filtering-firewalls/\">Docker Docs: Packet filtering and firewalls</a></li>\n<li><a href=\"https://docs.docker.com/reference/compose-file/services/#mem_limit\">Docker Compose services: <code>mem_limit</code></a></li>\n<li><a href=\"https://redis.io/docs/latest/operate/oss_and_stack/management/security/\">Redis security</a></li>\n<li><a href=\"https://redis.io/docs/latest/operate/oss_and_stack/management/persistence/\">Redis persistence</a></li>\n<li><a href=\"https://hub.docker.com/_/redis/\">Redis Official Image</a></li>\n</ul>\n","date_published":"2026-08-01T00:00:00.000Z","date_modified":"2026-08-01T00:00:00.000Z","tags":["Redis","Docker","Docker Compose","Ubuntu","Linux","Operations"],"language":"en"},{"id":"https://www.lihuanyu.com/posts/2026/%E7%94%A8-GitHub-Actions-%E6%8C%81%E7%BB%AD%E9%83%A8%E7%BD%B2%E5%BE%AE%E4%BF%A1%E5%B0%8F%E7%A8%8B%E5%BA%8F/","url":"https://www.lihuanyu.com/posts/2026/%E7%94%A8-GitHub-Actions-%E6%8C%81%E7%BB%AD%E9%83%A8%E7%BD%B2%E5%BE%AE%E4%BF%A1%E5%B0%8F%E7%A8%8B%E5%BA%8F/","title":"如何用 GitHub Actions 持续部署微信和支付宝小程序","summary":"使用微信官方 miniprogram-ci 和支付宝官方 minidev，把小程序构建、检查与版本上传接入 GitHub Actions。方案不依赖具体框架，原生项目和跨端框架都能使用。","content_html":"<p>微信官方的 <code>miniprogram-ci</code> 和支付宝官方的 <code>minidev</code> 都可以在命令行里上传小程序代码。把它们接入 GitHub Actions 后，主分支每次更新都能自动构建，并分别上传微信开发版本和支付宝体验版。</p>\n<p>下面这套配置不依赖具体框架。原生小程序可以用，Taro、uni-app、MPX 等跨端项目也可以用；只要最后能得到平台认识的构建目录即可。</p>\n<h2>部署流程</h2>\n<p>这套流程只关心构建产物，不关心源码使用什么框架：</p>\n<pre><code class=\"language-text\">git push\n  -&gt; 安装依赖\n  -&gt; 测试和构建\n  -&gt; 微信：miniprogram-ci 上传开发版本\n  -&gt; 支付宝：minidev 上传体验版\n</code></pre>\n<p>原生项目可以直接上传源码目录。跨端项目需要先执行各平台的构建命令，再上传对应的编译目录。测试和构建可以共用一个 job，微信、支付宝上传则适合拆成两个并行 job，互不影响。</p>\n<p>常见差异只有两项：</p>\n<div class=\"table-scroll\"><table>\n<thead>\n<tr>\n<th>项目类型</th>\n<th>构建命令</th>\n<th>上传目录</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>原生小程序</td>\n<td>可以省略</td>\n<td>平台工程配置文件所在目录</td>\n</tr>\n<tr>\n<td>Taro</td>\n<td>项目中的平台构建命令</td>\n<td>通常在 <code>dist</code> 下</td>\n</tr>\n<tr>\n<td>uni-app</td>\n<td>项目中的平台构建命令</td>\n<td>以项目配置的输出目录为准</td>\n</tr>\n<tr>\n<td>MPX</td>\n<td>项目中的跨端构建命令</td>\n<td>例如 <code>dist/wx</code>、<code>dist/ali</code></td>\n</tr>\n</tbody>\n</table>\n</div><p>部署脚本不需要理解这些框架。微信上传目录里要有 <code>project.config.json</code>，支付宝上传目录里要有 <code>mini.project.json</code>。</p>\n<h2>安装官方上传工具</h2>\n<p>在项目中安装 <a href=\"https://developers.weixin.qq.com/miniprogram/dev/devtools/ci.html\"><code>miniprogram-ci</code></a> 和 <a href=\"https://opendocs.alipay.com/mini/02q17h\"><code>minidev</code></a>：</p>\n<pre><code class=\"language-bash\">npm install --save-dev --save-exact miniprogram-ci minidev\n</code></pre>\n<p>建议锁定明确版本，并提交 lockfile。持续集成环境应该安装已经审核过的依赖解析结果，不要在部署时临时升级依赖。</p>\n<p>项目还需要相应的构建命令。例如：</p>\n<pre><code class=\"language-json\">{\n  &quot;scripts&quot;: {\n    &quot;test&quot;: &quot;your_test_command_here&quot;,\n    &quot;build:miniprograms&quot;: &quot;your_cross_platform_build_command_here&quot;\n  }\n}\n</code></pre>\n<p>只部署一个平台时，构建命令当然也可以只构建一个平台。原生小程序不需要编译时，可以删除构建脚本和 workflow 里的对应步骤。</p>\n<h2>生成微信代码上传密钥</h2>\n<p>登录微信公众平台，进入 <strong>开发管理 -&gt; 开发设置 -&gt; 小程序代码上传</strong>，生成代码上传密钥。</p>\n<p>下载得到的文件是 PEM 私钥：</p>\n<pre><code class=\"language-text\">-----BEGIN PRIVATE KEY-----\nprivate_key_content_here\n-----END PRIVATE KEY-----\n</code></pre>\n<p>密钥拥有预览和上传代码的权限。不要把它放进仓库，也不要输出到 Actions 日志。</p>\n<p>同一页面还可以配置 IP 白名单。普通 GitHub-hosted runner 没有固定出口 IP，本教程使用 <code>ubuntu-latest</code>，因此需要关闭微信代码上传 IP 白名单。需要保留白名单时，请改用具有固定出口 IP 的 runner。</p>\n<p>这是安全性与维护成本之间的选择。关闭白名单后，私钥成为主要凭证，更应该限制它的使用范围。</p>\n<h2>生成支付宝开发工具密钥</h2>\n<p>支付宝这里最容易卡住的不是密钥本身，而是入口藏得有点深。</p>\n<p>登录支付宝开放平台并进入控制台后，不要在某个小程序的开发设置里反复翻。正确路径在页面右上角：</p>\n<ol>\n<li>点击右上角头像</li>\n<li>进入 <strong>账户中心</strong></li>\n<li>进入 <strong>密钥管理</strong></li>\n<li>选择 <strong>开发工具密钥</strong></li>\n<li>生成身份密钥并下载 <code>config.json</code></li>\n</ol>\n<p>尤其是 <strong>账户中心</strong> 这一步，不点头像很难找到。知道路径以后不过几次点击，不知道时却足够在控制台里绕上几圈。</p>\n<p>下载得到的是完整的工具身份文件，大致如下：</p>\n<pre><code class=\"language-json\">{\n  &quot;alipay&quot;: {\n    &quot;authentication&quot;: {\n      &quot;toolId&quot;: &quot;...&quot;,\n      &quot;privateKey&quot;: &quot;-----BEGIN PRIVATE KEY-----...&quot;\n    }\n  }\n}\n</code></pre>\n<p>后面配置 <code>ALIPAY_IDENTITY_KEY</code> 时，要保存这个 <code>config.json</code> 的<strong>完整内容</strong>，不能只复制 <code>privateKey</code>。它也不是支付接口使用的应用私钥，两种密钥不要混在一起。</p>\n<p>同一个支付宝账号同时只有一份开发工具密钥有效。重新生成密钥，或者在其他电脑执行 <code>minidev login</code>，都会让旧密钥失效。CI 使用固定密钥文件，正是为了避免每台 runner 都扫码登录。</p>\n<p>顺手说一句，支付宝小程序目前本身不收费，对个人项目很友好。流量确实不算多，后台也谈不上好用，还有一套“小程序分”常常打得让人摸不着头脑。不过免费平台肯给上传接口，也就没什么可抱怨的了。</p>\n<h2>配置 GitHub Environment</h2>\n<p>打开仓库的 <strong>Settings -&gt; Environments</strong>，分别创建 <code>wechat-upload</code> 和 <code>alipay-upload</code> 两个 Environment。只部署一个平台时，创建对应的一个即可。</p>\n<p>在 <code>wechat-upload</code> 中添加：</p>\n<ul>\n<li><strong>Variable <code>WECHAT_APPID</code></strong>: 微信小程序 AppID</li>\n<li><strong>Secret <code>WECHAT_PRIVATE_KEY</code></strong>: 完整的 PEM 私钥内容</li>\n</ul>\n<p>在 <code>alipay-upload</code> 中添加：</p>\n<ul>\n<li><strong>Variable <code>ALIPAY_APP_ID</code></strong>: 支付宝小程序 AppID</li>\n<li><strong>Secret <code>ALIPAY_IDENTITY_KEY</code></strong>: 下载得到的 <code>config.json</code> 完整内容</li>\n</ul>\n<p>Environment 不是服务器。它用于管理部署变量、Secrets 和审批规则。上传任务声明对应的 <code>environment</code> 后，才能读取其中的值。</p>\n<p>如果仓库不需要环境隔离，也可以使用 Repository Variables 和 Repository Secrets。两种方式都能运行，Environment 的权限边界更清楚。</p>\n<h2>封装仓库内的上传 Action</h2>\n<p>把上传逻辑放进仓库内的 Composite Action，可以让 workflow 保持简短，也便于补充参数校验和测试。</p>\n<h3>微信 Action</h3>\n<p>创建目录：</p>\n<pre><code class=\"language-text\">.github/actions/wechat-upload/\n  action.yml\n  upload.cjs\n</code></pre>\n<p><code>action.yml</code> 声明 AppID、私钥和构建目录：</p>\n<pre><code class=\"language-yaml\">name: WeChat Mini Program Upload\ndescription: Upload a built WeChat Mini Program with miniprogram-ci.\n\ninputs:\n  appid:\n    description: WeChat Mini Program AppID.\n    required: true\n  private-key:\n    description: Content of the code upload private key.\n    required: true\n  project-path:\n    description: Directory containing project.config.json.\n    required: true\n\nruns:\n  using: composite\n  steps:\n    - shell: bash\n      env:\n        INPUT_APPID: ${{ inputs.appid }}\n        INPUT_PRIVATE_KEY: ${{ inputs.private-key }}\n        INPUT_PROJECT_PATH: ${{ inputs.project-path }}\n      run: node &quot;$GITHUB_ACTION_PATH/upload.cjs&quot;\n</code></pre>\n<p>Action 通过当前 step 的环境变量读取 Secret，不把私钥放进命令参数。</p>\n<p><code>upload.cjs</code> 调用微信官方接口：</p>\n<pre><code class=\"language-js\">const fs = require('node:fs')\nconst os = require('node:os')\nconst path = require('node:path')\nconst ci = require('miniprogram-ci')\n\nconst root = process.cwd()\nconst projectPath = path.resolve(process.env.INPUT_PROJECT_PATH)\nconst { version } = require(path.join(root, 'package.json'))\nconst tempDir = fs.mkdtempSync(path.join(os.tmpdir(), 'wechat-ci-'))\nconst keyPath = path.join(tempDir, 'private.key')\nconst privateKey = process.env.INPUT_PRIVATE_KEY.replace(/\\\\n/g, '\\n')\n\nfs.writeFileSync(keyPath, privateKey, { mode: 0o600 })\n\nconst project = new ci.Project({\n  appid: process.env.INPUT_APPID,\n  type: 'miniProgram',\n  projectPath,\n  privateKeyPath: keyPath\n})\n</code></pre>\n<p>创建项目对象后执行上传，并在结束时删除临时密钥：</p>\n<pre><code class=\"language-js\">;(async () =&gt; {\n  try {\n    await ci.upload({\n      project,\n      version,\n      desc: `Actions #${process.env.GITHUB_RUN_NUMBER} ` +\n        `(${process.env.GITHUB_SHA.slice(0, 7)})`,\n      robot: 1,\n      setting: {},\n      onProgressUpdate: console.log\n    })\n  } finally {\n    fs.rmSync(tempDir, { recursive: true, force: true })\n  }\n})().catch((error) =&gt; {\n  console.error(error)\n  process.exitCode = 1\n})\n</code></pre>\n<p>版本号直接读取 <code>package.json</code> 的 <code>version</code>。GitHub run number 和 commit SHA 放进版本描述，用来区分同一产品版本的不同构建。</p>\n<p>实际项目还可以增加这些校验：</p>\n<ul>\n<li>检查 <code>project.config.json</code> 是否存在</li>\n<li>检查配置文件中的 AppID 是否与 Action 输入一致</li>\n<li>限制 CI 机器人编号为 <code>1</code> 到 <code>30</code></li>\n<li>把版本号和描述压成单行</li>\n</ul>\n<p>上传失败应该直接结束任务。不要捕获错误后继续返回成功状态。</p>\n<h3>支付宝 Action</h3>\n<p>支付宝 Action 的结构相同，只是输入换成 AppID、身份密钥和支付宝构建目录：</p>\n<pre><code class=\"language-text\">.github/actions/alipay-upload/\n  action.yml\n  upload.cjs\n</code></pre>\n<p><code>action.yml</code> 把输入转成当前 step 的环境变量：</p>\n<pre><code class=\"language-yaml\">name: Alipay Mini Program Upload\ndescription: Upload a built Alipay Mini Program with minidev.\n\ninputs:\n  appid:\n    required: true\n  identity-key:\n    required: true\n  project-path:\n    required: true\n\nruns:\n  using: composite\n  steps:\n    - shell: bash\n      env:\n        INPUT_APPID: ${{ inputs.appid }}\n        INPUT_IDENTITY_KEY: ${{ inputs.identity-key }}\n        INPUT_PROJECT_PATH: ${{ inputs.project-path }}\n      run: node &quot;$GITHUB_ACTION_PATH/upload.cjs&quot;\n</code></pre>\n<p><code>upload.cjs</code> 把 <code>ALIPAY_IDENTITY_KEY</code> 的完整内容写进临时文件，再交给官方 <code>minidev</code>：</p>\n<pre><code class=\"language-js\">const fs = require('node:fs')\nconst os = require('node:os')\nconst path = require('node:path')\nconst { minidev } = require('minidev')\n\nconst root = process.cwd()\nconst projectPath = path.resolve(process.env.INPUT_PROJECT_PATH)\nconst { version } = require(path.join(root, 'package.json'))\nconst tempDir = fs.mkdtempSync(path.join(os.tmpdir(), 'alipay-ci-'))\nconst identityKeyPath = path.join(tempDir, 'config.json')\n\nfs.writeFileSync(identityKeyPath, process.env.INPUT_IDENTITY_KEY, { mode: 0o600 })\n\n;(async () =&gt; {\n  try {\n    await minidev.upload({\n      appId: process.env.INPUT_APPID,\n      clientType: 'alipay',\n      project: projectPath,\n      identityKeyPath,\n      version,\n      versionDescription: `Actions #${process.env.GITHUB_RUN_NUMBER}`,\n      experience: true\n    })\n  } finally {\n    fs.rmSync(tempDir, { recursive: true, force: true })\n  }\n})().catch((error) =&gt; {\n  console.error(error)\n  process.exitCode = 1\n})\n</code></pre>\n<p><code>experience: true</code> 会在上传成功后把该版本设为支付宝体验版。这个账号需要有对应权限，否则可以先去掉该参数，只上传版本。</p>\n<h2>编写部署 workflow</h2>\n<p>创建 <code>.github/workflows/deploy-miniprograms.yml</code>。工作流只响应主分支推送和手动触发，不响应 Pull Request（PR）。先做一次测试和双端构建，再把产物交给两个上传 job：</p>\n<pre><code class=\"language-yaml\">name: Deploy Mini Programs\n\non:\n  push:\n    branches:\n      - main\n  workflow_dispatch:\n\npermissions:\n  contents: read\n\njobs:\n  validate:\n    runs-on: ubuntu-latest\n    timeout-minutes: 20\n    steps:\n      - uses: actions/checkout@v7\n\n      - uses: actions/setup-node@v6\n        with:\n          node-version: lts/*\n          cache: npm\n\n      - run: npm ci\n      - run: npm test --if-present\n      - run: npm run build:miniprograms\n\n      - uses: actions/upload-artifact@v7\n        with:\n          name: wechat-miniprogram\n          path: dist/wx\n\n      - uses: actions/upload-artifact@v7\n        with:\n          name: alipay-miniprogram\n          path: dist/ali\n\n  upload-wechat:\n    needs: validate\n    runs-on: ubuntu-latest\n    environment: wechat-upload\n    concurrency:\n      group: wechat-upload\n      cancel-in-progress: false\n    steps:\n      - uses: actions/checkout@v7\n      - uses: actions/setup-node@v6\n        with:\n          node-version: lts/*\n          cache: npm\n      - run: npm ci\n\n      - uses: actions/download-artifact@v8\n        with:\n          name: wechat-miniprogram\n          path: dist/wx\n\n      - name: Upload WeChat development version\n        uses: ./.github/actions/wechat-upload\n        with:\n          appid: ${{ vars.WECHAT_APPID }}\n          private-key: ${{ secrets.WECHAT_PRIVATE_KEY }}\n          project-path: dist/wx\n\n  upload-alipay:\n    needs: validate\n    runs-on: ubuntu-latest\n    environment: alipay-upload\n    concurrency:\n      group: alipay-upload\n      cancel-in-progress: false\n    steps:\n      - uses: actions/checkout@v7\n      - uses: actions/setup-node@v6\n        with:\n          node-version: lts/*\n          cache: npm\n      - run: npm ci\n\n      - uses: actions/download-artifact@v8\n        with:\n          name: alipay-miniprogram\n          path: dist/ali\n\n      - name: Upload Alipay experience version\n        uses: ./.github/actions/alipay-upload\n        with:\n          appid: ${{ vars.ALIPAY_APP_ID }}\n          identity-key: ${{ secrets.ALIPAY_IDENTITY_KEY }}\n          project-path: dist/ali\n</code></pre>\n<p>把构建命令以及 <code>dist/wx</code>、<code>dist/ali</code> 换成项目自己的配置。使用 pnpm 或 Yarn 时，也只需要替换安装、缓存和脚本命令。</p>\n<p>两个上传 job 各自使用独立的 Environment 和并发组。连续推送时，同一平台的上传会排队，微信与支付宝之间则可以并行。</p>\n<h2>验证部署</h2>\n<p>提交 workflow、Action 和上传脚本后，推送到 <code>main</code>：</p>\n<pre><code class=\"language-bash\">git add .github package.json package-lock.json\ngit commit -m &quot;feat: deploy Mini Programs with GitHub Actions&quot;\ngit push origin main\n</code></pre>\n<p>打开仓库的 <strong>Actions</strong> 页面，依次确认：</p>\n<ol>\n<li>依赖安装成功</li>\n<li>测试和双端构建成功</li>\n<li>微信上传步骤显示 <code>miniprogram-ci</code> 编译与上传进度</li>\n<li>支付宝上传步骤显示 <code>minidev</code> 构建与上传结果</li>\n<li>微信后台出现新的开发版本，支付宝后台出现新的体验版</li>\n</ol>\n<p>微信上传提示 IP 不在白名单时，检查微信公众平台的 IP 白名单设置。支付宝提示身份校验失败时，先确认 <code>ALIPAY_IDENTITY_KEY</code> 保存的是完整 <code>config.json</code>，再检查开发工具密钥是否被重新生成过。</p>\n<h2>部署边界</h2>\n<p><code>miniprogram-ci.upload()</code> 对应微信开发者工具里的“上传”，只生成微信后台开发版本。<code>minidev.upload()</code> 会上传支付宝版本；示例中的 <code>experience: true</code> 还会把它设为体验版。两者都没有自动提审或正式发布。</p>\n<p>开发版本上传适合自动执行。提审和发布会直接影响用户，还涉及版本说明、隐私能力和审核材料，保留人工确认更稳妥。</p>\n<p>一条可靠的流水线不必包办所有事情。它应该把重复劳动拿走，把真正需要判断的地方留给人。</p>\n<h2>参考资料</h2>\n<ul>\n<li><a href=\"https://developers.weixin.qq.com/miniprogram/dev/devtools/ci.html\">微信小程序 CI 官方文档</a></li>\n<li><a href=\"https://opendocs.alipay.com/mini/02q17h\">支付宝小程序 CLI 官方文档</a></li>\n<li><a href=\"https://opendocs.alipay.com/mini/02q29w\">支付宝开发工具密钥</a></li>\n<li><a href=\"https://docs.github.com/actions/deployment/targeting-different-environments/managing-environments-for-deployment\">GitHub Actions Environments</a></li>\n<li><a href=\"https://docs.github.com/actions/security-guides/using-secrets-in-github-actions\">GitHub Actions Secrets</a></li>\n</ul>\n","date_published":"2026-07-11T00:00:00.000Z","tags":["GitHub Actions","微信小程序","支付宝小程序","CI CD","miniprogram-ci","minidev","自动化部署"],"language":"zh"},{"id":"https://www.lihuanyu.com/en/posts/2026/locks-are-boundaries-not-syntax/","url":"https://www.lihuanyu.com/en/posts/2026/locks-are-boundaries-not-syntax/","title":"Locks Are Not Syntax, They Are Boundaries","summary":"A practical explanation of what locks are for, when to consider them, how optimistic and pessimistic locking differ, and why local locks are not enough in distributed systems.","content_html":"<p>Imagine a paper sign-out sheet sitting by the office door.</p>\n<p>Who borrowed the projector, who took the meeting room key, who picked up the last box of printer paper: everyone has to write it down. When only one person is writing, everything is fine. Name, time, item. Done.</p>\n<p>The trouble begins when two people try to write on the same line.</p>\n<p>One person has written a name but not the time yet. Another person crosses out that line and writes their own record. A third person sees that the sheet still says one box of printer paper is available and takes it away. In the end, all three people feel they did nothing wrong, but the ledger is already a mess.</p>\n<p>That is the problem locks are meant to solve.</p>\n<p>A lock is not some advanced piece of language syntax. It is not there to make code look more “senior”. It is first of all a boundary: when many people may modify the same thing at the same time, the system needs a way to decide who goes first, who waits, how failure is handled, and what final result counts as correct.</p>\n<p>This idea is worth revisiting now that many people write code with AI.</p>\n<p>AI can quickly produce APIs, forms, validation, database access, and tests. It will also become better at deciding where a lock is needed, where a transaction is needed, and where idempotency is needed. But developers still need to understand why it made those choices.</p>\n<p>If it adds a lock, you need to know whether it locked the right resource.</p>\n<p>If it does not add a lock, you need to know whether this place really does not need one.</p>\n<p>If it uses a distributed lock, you need to know whether that lock is truly visible to all instances, or whether it only locked the door of one small room.</p>\n<p><img src=\"/assets/posts/2026/locks-ai-architecture/01-human-architecture-judgment.jpg\" alt=\"AI can generate code, but people still decide system boundaries\"></p>\n<p><a href=\"/posts/locks-are-boundaries-not-syntax/\">Chinese version of this article</a></p>\n<h2>Locks Solve Disorder, Not Slowness</h2>\n<p>When people first learn about locks, they often connect them with performance.</p>\n<p>That is understandable. Locks often make programs slower. One person gets the lock, others wait. The larger the lock scope, the more people wait. The longer the lock is held, the more throughput falls.</p>\n<p>But the first problem a lock solves is not speed. It solves correctness.</p>\n<p>Without a lock, or without another form of concurrency control, a program may run very fast. It may simply write bad data very fast.</p>\n<p>In the paper sign-out sheet example, nobody has to wait. That is certainly the fastest version. The cost is that nobody knows who actually borrowed the projector. Meeting room reservations are similar. The system shows that 2 p.m. is available. Two teams click reserve at the same time. If there is no constraint in between, the same room may be assigned to both teams.</p>\n<p>This is not just a display bug.</p>\n<p>In a real system, that sheet may be an inventory table, an account balance row, an order status row, a coupon redemption record, or a task queue table. These records can be read and modified. Once multiple requests modify the same data at the same time, the system needs an answer.</p>\n<p>Who modifies first?</p>\n<p>Who modifies later?</p>\n<p>If the later request discovers that the data has changed, should it retry, fail, or reload?</p>\n<p>A lock is one of the most common answers.</p>\n<h2>When to Think About Locks</h2>\n<p>Do not add a lock the moment you see the word “concurrency”.</p>\n<p>A steadier way is to ask three questions.</p>\n<p>First, can multiple execution units happen at the same time?</p>\n<p>An execution unit is not only a Java thread. It can be two HTTP requests, two background jobs, two service instances, two message consumers, or two users clicking the same button in the same second.</p>\n<p>Second, do they access the same resource?</p>\n<p>The same product, the same account, the same meeting room, the same coupon, the same task record: all of these count.</p>\n<p>Third, is there a modification?</p>\n<p>Pure reads are usually easier. The trouble starts when at least one side writes. Two requests checking whether a meeting room is free are not the problem. Two requests both changing the same room to “reserved” are.</p>\n<p>If all three answers are yes, you should consider concurrency control.</p>\n<p>Consider concurrency control, not immediately hand-write a lock.</p>\n<p>Some cases are suited to locks. Some are better handled by a database unique index. Some only need a conditional update. Some should be turned into queue-based serial processing. A lock is a tool. Not every door needs to become a vault door.</p>\n<h2>What Is Actually Being Locked</h2>\n<p>Beginners often think a lock locks a piece of code.</p>\n<p>For example, in Java, <code>synchronized</code> looks as if it locks a method or a code block. That description is convenient, but it can mislead.</p>\n<p>The real question is: what resource is this lock protecting?</p>\n<p>If it protects inventory, then decrementing the inventory of the same product should be mutually exclusive. Different products usually do not need to wait for each other.</p>\n<p>If it protects an account balance, then deductions and deposits on the same account need care. Different accounts usually should not line up in one long queue.</p>\n<p>If it protects a meeting room, then Room A and Room B can be reserved independently. Room B should not be blocked just because Room A is locked.</p>\n<p>Lock granularity matters.</p>\n<p>If the lock is too small, it does not protect the problem. You meant to protect one account balance, but you only locked a temporary object. Another service instance can still update the database.</p>\n<p>If the lock is too large, the system gets jammed. If one product’s inventory update blocks every product in the shop, that is like cutting power to an entire building so two people do not fight over one pen.</p>\n<p>So more locks are not automatically safer. Larger locks are not automatically safer either.</p>\n<p>A good lock protects exactly the resource that can be written incorrectly.</p>\n<h2>The Smallest Example: <code>count++</code></h2>\n<p>Technical articles often use <code>count++</code> to explain locks. It is old-fashioned, but it works.</p>\n<pre><code class=\"language-java\">count++;\n</code></pre>\n<p>That line looks like one action. In reality, it can be understood as three steps:</p>\n<pre><code class=\"language-text\">read count\ncalculate count + 1\nwrite count back\n</code></pre>\n<p>If two threads both read <code>count = 10</code>, each calculates 11, and both write 11 back, the final result is 11, not 12.</p>\n<p>Inventory deduction, balance changes, and order status transitions are business versions of this same problem.</p>\n<p>The difference is that if <code>count++</code> is wrong, maybe one statistic is off by one. If a balance is wrong, it is no longer just a number being off by one.</p>\n<p>The problem is not that this line of code is complicated.</p>\n<p>The problem is that it pretends only one person exists in the world.</p>\n<h2>Pessimistic Locking: Close the Door First</h2>\n<p>Pessimistic locking is a plain idea: I believe conflict may happen, so I lock the resource first. When I finish, others can come in.</p>\n<p>It is like entering a small room that can hold only one person. You close the door after entering. Not because someone will definitely rush in, but because if it happens, the result is ugly.</p>\n<p>Java’s <code>synchronized</code>, <code>ReentrantLock</code>, and a database’s <code>SELECT ... FOR UPDATE</code> can all be understood through this idea.</p>\n<p>A common database form looks like this:</p>\n<pre><code class=\"language-sql\">SELECT balance\nFROM account\nWHERE id = ?\nFOR UPDATE;\n\nUPDATE account\nSET balance = balance - ?\nWHERE id = ?;\n</code></pre>\n<p>This usually runs inside one transaction. Before the transaction commits, that account row is locked. Other transactions that want to modify the same row have to wait.</p>\n<p>Where does pessimistic locking fit?</p>\n<p>It fits places with higher conflict, data that must not be wrong, and failures that are not easy to retry. Account deductions, critical order status transitions, and some strongly constrained resource reservations are typical examples.</p>\n<p>Its cost is also direct.</p>\n<p>Others have to wait. Holding a lock too long hurts throughput. Taking multiple locks in a messy order can cause deadlocks. A pessimistic lock is like temporarily closing a road. Useful when needed; painful when overused.</p>\n<p>One practical rule is: do not do slow work while holding a lock.</p>\n<p>In particular, avoid calling remote services while holding a database row lock. That is like locking everyone outside while you make a phone call inside the room. If the call lasts five minutes, the people outside will move from polite patience to questioning their life choices.</p>\n<h2>Optimistic Locking: Go First, Check the Ticket Later</h2>\n<p>Optimistic locking takes the opposite attitude.</p>\n<p>It does not stop people in advance. Everyone works first. When a modification is submitted, the system checks whether the data changed in the meantime.</p>\n<p>A common approach is a <code>version</code> column.</p>\n<p>When updating inventory, include the old version:</p>\n<pre><code class=\"language-sql\">UPDATE product\nSET stock = stock - 1,\n    version = version + 1\nWHERE id = ?\n  AND stock &gt; 0\n  AND version = ?;\n</code></pre>\n<p>If the affected row count is 1, the update succeeded.</p>\n<p>If it is 0, either inventory is insufficient or the row was already changed by someone else. The next step may be retrying, or it may be returning failure.</p>\n<p>Optimistic locking fits low-conflict, read-heavy cases where failure and retry are acceptable: editing an article, updating configuration, or ordinary inventory deduction.</p>\n<p>Its trouble is failure handling.</p>\n<p>Many implementations only write the success path. What happens if the update fails? Retry how many times? What does the user see? Can retries make the system even busier?</p>\n<p>Those questions cannot be skipped.</p>\n<p>Optimistic locking is not “no lock”. It turns waiting into checking, and turns queuing into failure handling.</p>\n<p>There is also a common term called ABA. It means data changed from A to B and then back to A. It looks unchanged, but something happened in the middle. In many business systems, an increasing version number is enough to avoid this. There is no need to start from the lowest-level details.</p>\n<h2>How to Choose</h2>\n<p>Pessimistic locking means closing the door first, then doing the work.</p>\n<p>Optimistic locking means doing the work first, then checking the ticket.</p>\n<p>Neither is more advanced. Comparing technologies without a scenario usually turns into a fight between terms. The terms may fight loudly. The production system does not care.</p>\n<p>A rough guide:</p>\n<div class=\"table-scroll\"><table>\n<thead>\n<tr>\n<th>Question</th>\n<th>More likely choice</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>High conflict and retry is expensive</td>\n<td>Pessimistic locking</td>\n</tr>\n<tr>\n<td>Low conflict and retry is acceptable</td>\n<td>Optimistic locking</td>\n</tr>\n<tr>\n<td>Data must not be wrong, such as balances</td>\n<td>Transactions, row locks, ledgers, idempotency together</td>\n</tr>\n<tr>\n<td>Only need to prevent duplicate creation</td>\n<td>Unique indexes and idempotency keys are often more direct</td>\n</tr>\n<tr>\n<td>Multi-instance deployment</td>\n<td>Local locks are usually not enough; use external constraints</td>\n</tr>\n</tbody>\n</table>\n</div><p>In engineering, some answers really are “it depends”. What matters is not memorizing one fixed choice, but laying out the conditions.</p>\n<p>How likely is conflict?</p>\n<p>How bad is a wrong result?</p>\n<p>Can the operation be retried?</p>\n<p>Is the system single-process or multi-instance?</p>\n<p>Where is the most reliable place to enforce the constraint?</p>\n<p>Once these conditions are clear, the solution usually has a direction.</p>\n<p><img src=\"/assets/posts/2026/locks-ai-architecture/02-shared-data-collision.jpg\" alt=\"Multiple requests meet at one shared data gate; the real question is whether the shared resource can be written incorrectly\"></p>\n<h2>Meeting Rooms, Inventory, and Balances</h2>\n<p>The easiest everyday example is meeting room reservation.</p>\n<p>The same meeting room at the same time can only be reserved by one team. The solution does not have to be a hand-written lock. A database unique constraint can do the job: room ID plus time range must not duplicate. Whoever writes first succeeds; the later write fails.</p>\n<p>This is like saying the sign-out sheet cannot contain two rows with the same record number.</p>\n<p>Inventory deduction is similar.</p>\n<p>The easiest wrong version looks like this:</p>\n<pre><code class=\"language-text\">check inventory\nif inventory is greater than 0\ndeduct inventory\n</code></pre>\n<p>Single-user testing passes. Under concurrency, two requests may both read <code>stock = 1</code>, both believe they can buy, and the product is oversold.</p>\n<p>A more direct approach is to combine the check and the modification into one SQL statement:</p>\n<pre><code class=\"language-sql\">UPDATE product\nSET stock = stock - 1\nWHERE id = ?\n  AND stock &gt; 0;\n</code></pre>\n<p>If the affected row count is 1, the deduction succeeded.</p>\n<p>If it is 0, inventory was insufficient.</p>\n<p>This SQL is plain, almost boring, but useful. It puts the constraint “stock must be greater than 0” inside the database operation and avoids someone slipping in between “check” and “update”.</p>\n<p>Balances are more sensitive.</p>\n<p>Overselling inventory is ugly, but it can still be handled by restocking, refunding, and apologizing. If balances are wrong, the system loses a piece of trust. The scariest thing in an accounting-like system is not ugly code. It is books that do not reconcile.</p>\n<p>For balance changes, thinking only about “adding a lock” is not enough. At least transactions, ledgers, idempotency, and unique indexes need to be considered.</p>\n<p>A transaction makes the balance update and ledger insert succeed or fail together.</p>\n<p>A ledger makes every change traceable.</p>\n<p>Idempotency ensures the same deduction request is not executed twice because of a retry.</p>\n<p>A unique index stops duplicate request IDs at the database boundary.</p>\n<p>This shows an important point: concurrency control is not only locks.</p>\n<p>A lock is one tool among several.</p>\n<h2>Duplicate Submission Is Not Just a Lock Problem</h2>\n<p>Users double-click buttons, browsers resend requests, clients retry after timeouts, and message queues redeliver messages. All of these can cause duplicate submission.</p>\n<p>Take payment as an example.</p>\n<p>A user clicks pay. The network stalls. They click again. Or the frontend sent only one request, but the gateway retried. If the system creates a new payment order for every request, the rest becomes lively in the worst possible way.</p>\n<p>Many people first think of adding a lock.</p>\n<p>Sometimes that works. It is not always the best answer.</p>\n<p>The key point in duplicate submission is that the same business action should be processed only once. This is closer to idempotency.</p>\n<p>The word idempotency sounds a little mathematical. In business systems, it can be understood simply: the same voucher may arrive several times, but only the first one counts; later arrivals return the same result.</p>\n<p>Common approaches include:</p>\n<ul>\n<li>Disable the frontend button to improve experience, but do not treat it as backend protection.</li>\n<li>Use an <code>idempotency_key</code> or request ID on the backend.</li>\n<li>Add a database unique index to block duplicate request IDs.</li>\n<li>Return the first processing result for duplicate requests.</li>\n<li>Use a short Redis lock to absorb very short repeated clicks if needed.</li>\n</ul>\n<p>Locks prevent “entering at the same time”.</p>\n<p>Idempotency prevents “doing the same thing multiple times”.</p>\n<p>They are related, but they are not the same thing.</p>\n<h2>Repeated Jobs: The Door Is Not in the Same Room</h2>\n<p>In local development, a scheduled job usually runs in one process.</p>\n<p>Production changes that.</p>\n<p>Three instances may each have the same cron job. When the time comes, all three scan the database. You meant to process one batch of data, and it gets processed three times.</p>\n<p>Java’s <code>synchronized</code> does not help here.</p>\n<p>Each instance has its own JVM, memory, and locks. Machine A locks something; Machine B does not know.</p>\n<p>A local lock is like the door lock of one room. In a multi-instance deployment, there are several rooms, each with its own lock. You locked your own door, while the room next door can still modify the same remote ledger.</p>\n<p>This kind of situation needs a public gate that all instances can see.</p>\n<p>Common choices:</p>\n<ul>\n<li>A Redis distributed lock, allowing only one instance to execute at a time.</li>\n<li>A database task table claim, where whoever updates the state successfully owns the task.</li>\n<li>A message queue, splitting work into messages consumed under queue rules.</li>\n<li>A scheduler that guarantees single-instance execution for the job.</li>\n</ul>\n<p>The point is not to memorize a framework. The point is to know that local locks and distributed coordination are not the same thing.</p>\n<h2>A Distributed Lock Is Not a Luxury <code>synchronized</code></h2>\n<p>When multi-instance deployment comes up, Redis distributed locks often come to mind.</p>\n<p>They are useful, but do not think of them as a remote luxury version of <code>synchronized</code>.</p>\n<p>Distributed locks have to face real problems:</p>\n<ul>\n<li>How long should the lock live?</li>\n<li>What if business execution exceeds the lock expiration time?</li>\n<li>Can the lock be released if the service crashes?</li>\n<li>Is renewal needed?</li>\n<li>Can releasing the lock delete someone else’s newer lock?</li>\n</ul>\n<p>The basic requirement is that the lock has an expiration time, the lock value has a unique identity, and release first verifies that the current holder really owns the lock.</p>\n<p>Otherwise, a strange situation can happen.</p>\n<p>A acquires the lock. The business code runs too slowly, and the lock expires. B acquires a new lock and starts processing. A finally finishes and deletes B’s lock.</p>\n<p>This is not folklore. It is an ordinary awkwardness that real systems can hit.</p>\n<p>That is why in many cases, database unique indexes, conditional updates, or task-table claiming are steadier than a hand-written distributed lock.</p>\n<p>A distributed lock can be used. But you need to know what you are holding.</p>\n<h2>State Transitions: Draw the Roads First</h2>\n<p>An order may move from “pending payment” to “paid”, then to “shipped”. It may also be refunded, canceled, or closed.</p>\n<p>These states are not arbitrary.</p>\n<p>Pending payment can become paid or canceled. After shipping, an order cannot pretend to go back to pending payment. Refunds are not allowed at every moment either.</p>\n<p>Under concurrency, payment callbacks, user cancellations, admin shipping actions, and refund requests may arrive at the same time. If the final status becomes messy, reading logs starts to feel like archaeology.</p>\n<p>Locks solve only part of this problem.</p>\n<p>The deeper issue is state constraints.</p>\n<p>A common pattern is a conditional update:</p>\n<pre><code class=\"language-sql\">UPDATE orders\nSET status = 'PAID'\nWHERE id = ?\n  AND status = 'PENDING_PAYMENT';\n</code></pre>\n<p>Only when the order is still pending payment can it become paid. If the affected row count is 0, the status has already changed, and the current operation cannot pretend to have succeeded.</p>\n<p>This is the plain version of a state machine.</p>\n<p>A lot of bad code is too confident: <code>set status = xxx</code>, regardless of the old state.</p>\n<p>In state transitions, a lock is a door, and the state machine is a road map. If you install doors but never draw the roads, people still get lost.</p>\n<h2>Different Languages Use Different Door Locks</h2>\n<p>Java has <code>synchronized</code>, <code>ReentrantLock</code>, and <code>Atomic</code>.</p>\n<p>Go has <code>sync.Mutex</code>, <code>sync.RWMutex</code>, and often uses channels to turn shared modification into queue-like processing.</p>\n<p>Python has <code>threading.Lock</code>, and async code has <code>asyncio.Lock</code>.</p>\n<p>Node.js has an event-loop model on the main thread, but server-side systems still face concurrent requests, multiple processes, multiple instances, and database concurrency.</p>\n<p>Do not be fooled by “Node is single-threaded”.</p>\n<p>That sentence only means a piece of JavaScript code is not executed by two CPUs at the same time in the same thread. It does not stop two HTTP requests from updating the same database row, and it does not stop two machines from running the same job.</p>\n<p>Language-level locks mostly solve in-process problems.</p>\n<p>Business-system concurrency often happens between the database, cache, message queue, and multiple service instances.</p>\n<p>So the first question is not which language you use.</p>\n<p>The first question is: where is the shared data, where are the concurrent entrances, and who guarantees correctness?</p>\n<h2>Back to AI</h2>\n<p>AI will become better at writing code and planning solutions.</p>\n<p>That is a good thing.</p>\n<p>But developers cannot hand over basic concepts entirely. Whether a system breaks usually depends less on whether the code looks plausible and more on whether the constraints were placed in the right location.</p>\n<p>If AI adds <code>synchronized</code> to an API and the service is deployed in multiple instances, you need to know that this lock may only protect the current JVM.</p>\n<p>If AI writes “check first, then update” for inventory deduction, you need to know that concurrency can slip in between.</p>\n<p>If AI adds a Redis lock for duplicate submissions, you still need to check whether there is an idempotency key and a unique index.</p>\n<p>If AI directly overwrites an order status with <code>PAID</code>, you should ask: what was the old status, and is this transition legal?</p>\n<p>This does not mean everyone has to memorize JVM, AQS, JMM, or <code>volatile</code> details. Those things matter, but not every piece of business code needs to drill down to the center of the earth.</p>\n<p>The more common and more important questions are:</p>\n<ul>\n<li>Is there a shared resource?</li>\n<li>Is there concurrent modification?</li>\n<li>How bad is a wrong result?</li>\n<li>Can failure be retried?</li>\n<li>Is this single-process or multi-instance?</li>\n<li>Should the constraint live in code, the database, Redis, a queue, or the scheduler?</li>\n</ul>\n<p>Once these questions are clear, you can tell whether the solution AI gives you is reliable.</p>\n<p>If you cannot understand them, faster code generation can simply bring faster mistakes.</p>\n<p>That is not meant to be scary. It is ordinary engineering sense: when the car gets faster, road signs and brakes have to keep up.</p>\n<h2>Finally</h2>\n<p>Locks solve correctness problems when shared resources are modified concurrently.</p>\n<p>Pessimistic locking fits high-conflict cases where failure is hard to accept. Close the door first, then do the work.</p>\n<p>Optimistic locking fits lower-conflict cases where retry is acceptable. Do the work first, then check the ticket.</p>\n<p>In distributed systems, local locks are usually not enough. Database locks, Redis distributed locks, unique indexes, conditional updates, message queues, and task-table claiming may all be part of the answer.</p>\n<p>In real business code, do not start by asking, “Should I add a lock?”</p>\n<p>Ask first:</p>\n<ul>\n<li>Can this business flow be wrong under concurrency?</li>\n<li>How bad is the impact if it is wrong?</li>\n<li>Where is the best place to enforce the constraint?</li>\n<li>Which resource is the AI or framework solution actually locking?</li>\n</ul>\n<p>A lock is not an advanced concept.</p>\n<p>It simply reminds us that some doors in a system should not let everyone squeeze through at once.</p>\n<p>As for which door to close, how long to close it, and whether to use a key, an access card, or a sign-out sheet, people still have to think that through.</p>\n","date_published":"2026-07-10T00:00:00.000Z","tags":["Concurrency Control","Locks","Database","Backend","Architecture"],"language":"en"},{"id":"https://www.lihuanyu.com/posts/locks-are-boundaries-not-syntax/","url":"https://www.lihuanyu.com/posts/locks-are-boundaries-not-syntax/","title":"锁不是语法，是一种边界","summary":"从公共账本、会议室预约、库存和余额几个例子说起，理解锁到底解决什么问题，什么时候需要考虑锁，乐观锁和悲观锁怎么选，以及分布式场景里为什么不能只靠本机锁。","content_html":"<p>想象一张放在办公室门口的纸质登记表。</p>\n<p>谁借走了投影仪，谁拿了会议室钥匙，谁领了最后一盒打印纸，都要在上面写一笔。只有一个人登记时，一切都很正常。名字、时间、物品，写完就走。</p>\n<p>麻烦出现在两个人同时来写同一行。</p>\n<p>一个人刚写了名字，还没写时间；另一个人已经把那一行划掉，写成了自己的记录。再来一个人，看见表上还有一盒打印纸，就拿走了。最后三个人都觉得自己没错，账本却乱了。</p>\n<p>这就是锁要解决的问题。</p>\n<p>锁不是某种语言里的高级语法，也不是为了让代码看起来更像“资深工程师写的”。它首先是一条边界：当很多人可能同时修改同一份东西时，系统需要一种办法决定谁先来、谁等待、失败怎么处理，以及最后结果怎么算正确。</p>\n<p>现在用 AI 写代码，这个概念反而更值得重新讲一遍。</p>\n<p>AI 可以很快写出接口、表单、校验、数据库访问和测试用例。以后它也会越来越会自己判断哪里该加锁、哪里该用事务、哪里该做幂等。但开发者至少要看得懂它为什么这么做。</p>\n<p>它加了锁，你要知道锁住的是不是正确的资源。</p>\n<p>它没加锁，你要知道这里是不是本来就不需要。</p>\n<p>它用了分布式锁，你要知道这是不是一把真正大家都看得见的锁，还是只锁住了自己屋里的门。</p>\n<p><img src=\"/assets/posts/2026/locks-ai-architecture/01-human-architecture-judgment.jpg\" alt=\"AI 可以生成代码，人仍然要决定系统边界\"></p>\n<p><a href=\"/en/posts/2026/locks-are-boundaries-not-syntax/\">English version: Locks Are Not Syntax, They Are Boundaries</a></p>\n<h2>锁解决的不是慢，是乱</h2>\n<p>很多人第一次接触锁，会把它和性能放在一起想。</p>\n<p>这也正常。锁经常会让程序变慢。一个人拿到锁，其他人要等；锁的范围大了，等待的人就更多；锁持有久了，吞吐就掉下来。</p>\n<p>但锁最先解决的不是快慢问题，而是正确性问题。</p>\n<p>没有锁，或者没有其他并发控制手段，程序可能跑得很快。只是它会很快地把数据写乱。</p>\n<p>纸质登记表的例子里，所有人都不用等，当然最快。代价是最后没人知道投影仪到底被谁借走了。会议室预约也是一样。系统显示下午两点还有空，两个团队同时点预约。如果中间没有任何约束，最后同一间会议室可能被订给两拨人。</p>\n<p>这不是“页面显示错了”那么简单。</p>\n<p>真实业务里，这张登记表可能是库存表、账户余额表、订单状态表、优惠券领取记录、任务队列表。它们都可以被读，也可以被改。一旦多个请求同时修改同一份数据，系统就需要一个说法。</p>\n<p>谁先改？</p>\n<p>谁后改？</p>\n<p>后来的请求发现数据变了，是重试，失败，还是重新读取？</p>\n<p>锁就是这些说法里最常见的一种。</p>\n<h2>什么时候要想到锁</h2>\n<p>不用看到并发两个字就立刻加锁。</p>\n<p>更稳妥的办法，是先问三个问题。</p>\n<p>第一，会不会有多个执行单元同时发生？</p>\n<p>这里的执行单元不只是 Java 线程。它可以是两个 HTTP 请求，两个后台任务，两个服务实例，两个消息消费者，也可以是两个用户在同一秒点了同一个按钮。</p>\n<p>第二，它们会不会访问同一份资源？</p>\n<p>同一件商品、同一个账户、同一间会议室、同一张优惠券、同一条任务记录，都算。</p>\n<p>第三，里面有没有修改？</p>\n<p>只读通常问题不大。麻烦出在至少有一个人要改。两个请求同时查会议室是否空闲，没有关系；两个请求同时把同一间会议室改成“已预约”，就有关系。</p>\n<p>这三个问题如果都成立，就要考虑并发控制。</p>\n<p>注意，是考虑并发控制，不是立刻手写一把锁。</p>\n<p>有些场景用锁合适，有些场景用数据库唯一索引更直接，有些场景用条件更新就够，有些场景应该改成消息队列排队处理。锁是一件工具，不是所有门都要换成防盗门。</p>\n<h2>锁住的到底是什么</h2>\n<p>很多初学者会以为锁住的是一段代码。</p>\n<p>比如 Java 里写了 <code>synchronized</code>，看起来像是把某个方法或代码块锁住了。这个说法方便，但容易误导。</p>\n<p>真正要想的是：这把锁保护的资源是什么？</p>\n<p>如果保护的是库存，那同一个商品的库存扣减应该互斥，不同商品之间未必需要互相等待。</p>\n<p>如果保护的是账户余额，那同一个账户的扣款和充值要小心，不同账户之间通常不该排成一队。</p>\n<p>如果保护的是一间会议室，那 A 会议室和 B 会议室可以各自预约，不需要因为 A 被锁住，B 也不能用。</p>\n<p>锁的粒度很关键。</p>\n<p>锁太小，挡不住问题。该保护同一个账户余额，却只锁住了某个临时对象，另一个服务实例照样能改数据库。</p>\n<p>锁太大，系统就堵。只要有人改一个商品库存，全站所有商品都不能下单，这就像为了防止两个人抢一支笔，把整栋楼停电。</p>\n<p>所以锁不是越多越安全，也不是越大越安全。</p>\n<p>好锁要刚好保护那份会被写坏的资源。</p>\n<h2>最小的例子：<code>count++</code></h2>\n<p>技术文章里常用 <code>count++</code> 讲锁。它看起来老套，但确实好用。</p>\n<pre><code class=\"language-java\">count++;\n</code></pre>\n<p>这一行像一个动作，实际可以拆成三步：</p>\n<pre><code class=\"language-text\">读取 count\n计算 count + 1\n写回 count\n</code></pre>\n<p>如果两个线程同时读到 <code>count = 10</code>，它们各自算出 11，再各自写回 11，最后结果就是 11，不是 12。</p>\n<p>库存扣减、余额变更、订单状态流转，本质上都是这个问题的业务版本。</p>\n<p>只不过 <code>count++</code> 错了，可能只是统计数字少 1。余额错了，就不只是数字少 1 了。</p>\n<p>问题不在这一行代码有多复杂。</p>\n<p>问题在它假装世界上只有一个人在操作。</p>\n<h2>悲观锁：先关门，再办事</h2>\n<p>悲观锁的想法很朴素：我认为冲突可能发生，所以先把资源锁住，等我处理完，别人再来。</p>\n<p>像进一间只能容纳一个人的小房间。进去先关门，不是因为门外一定有人闯进来，而是这件事如果发生，会很难看。</p>\n<p>Java 里的 <code>synchronized</code>、<code>ReentrantLock</code>，数据库里的 <code>SELECT ... FOR UPDATE</code>，都可以按这个思路理解。</p>\n<p>数据库里常见的写法是：</p>\n<pre><code class=\"language-sql\">SELECT balance\nFROM account\nWHERE id = ?\nFOR UPDATE;\n\nUPDATE account\nSET balance = balance - ?\nWHERE id = ?;\n</code></pre>\n<p>这段 SQL 通常放在同一个事务里。事务提交前，这一行账户记录被锁住。别的事务如果也想改同一行，就要等。</p>\n<p>悲观锁适合什么场景？</p>\n<p>冲突比较高，数据不能错，失败后重试也不太合适。比如账户扣款、关键订单状态流转、某些强约束的资源占用。</p>\n<p>它的代价也很直白。</p>\n<p>别人要等。锁持有太久会影响吞吐。多把锁拿取顺序乱了，还可能死锁。悲观锁像临时封路，需要时很好用，封太多就堵车。</p>\n<p>一个很实际的经验是：锁里面少做慢事。</p>\n<p>尤其不要在拿着数据库行锁时去调远程接口。那就像把别人关在门外，自己在屋里打电话。电话一聊五分钟，门口的人会从礼貌等待变成怀疑人生。</p>\n<h2>乐观锁：先走，提交时验票</h2>\n<p>乐观锁的想法相反。</p>\n<p>它不提前拦人。大家先做自己的事，提交修改时检查一下：这份数据中途有没有被别人改过？</p>\n<p>常见做法是加 <code>version</code> 字段。</p>\n<p>更新库存时带上旧版本号：</p>\n<pre><code class=\"language-sql\">UPDATE product\nSET stock = stock - 1,\n    version = version + 1\nWHERE id = ?\n  AND stock &gt; 0\n  AND version = ?;\n</code></pre>\n<p>影响行数是 1，说明更新成功。</p>\n<p>影响行数是 0，说明库存不足，或者这条记录已经被别人改过。接下来可以重试，也可以返回失败。</p>\n<p>乐观锁适合冲突不高、读多写少、允许失败后重试的场景。比如编辑文章、更新配置、普通库存扣减。</p>\n<p>它的问题在失败处理。</p>\n<p>很多代码写乐观锁，只写成功路径。更新失败了怎么办？重试几次？用户看到什么？重试会不会把系统打得更忙？</p>\n<p>这些不能省。</p>\n<p>乐观锁不是“没有锁”。它只是把等待变成了检查，把排队变成了失败后的处理。</p>\n<p>还有一个常见词叫 ABA。意思是数据从 A 变成 B，又变回 A，表面看没变，其实中间发生过修改。多数业务里用递增版本号就能避开，不需要一上来钻到底层。</p>\n<h2>怎么选</h2>\n<p>悲观锁是先关门，再办事。</p>\n<p>乐观锁是先办事，最后验票。</p>\n<p>这两个没有谁更高级。脱离场景比较技术，通常会变成名词互殴。名词打得很热闹，线上系统不认。</p>\n<p>大致可以这样看：</p>\n<div class=\"table-scroll\"><table>\n<thead>\n<tr>\n<th>问题</th>\n<th>更可能的选择</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>冲突很多，失败重试成本高</td>\n<td>悲观锁</td>\n</tr>\n<tr>\n<td>冲突不多，允许用户重试或系统自动重试</td>\n<td>乐观锁</td>\n</tr>\n<tr>\n<td>数据绝不能错，比如余额</td>\n<td>事务、行锁、流水、幂等一起考虑</td>\n</tr>\n<tr>\n<td>只是防止重复创建</td>\n<td>唯一索引、幂等 key 往往更直接</td>\n</tr>\n<tr>\n<td>多实例部署</td>\n<td>本机锁通常不够，要考虑外部约束</td>\n</tr>\n</tbody>\n</table>\n</div><p>工程里有些答案确实只能说“不一定”。真正有用的不是背一个固定选择，而是把条件摆出来。</p>\n<p>冲突概率有多高。</p>\n<p>错误代价有多大。</p>\n<p>失败后能不能重试。</p>\n<p>系统是单进程还是多实例。</p>\n<p>约束到底放在哪里最可靠。</p>\n<p>这些条件说清楚，方案就不会差太远。</p>\n<p><img src=\"/assets/posts/2026/locks-ai-architecture/02-shared-data-collision.jpg\" alt=\"多路请求汇到同一份数据前，真正的问题是共享资源会不会被写坏\"></p>\n<h2>会议室、库存和余额</h2>\n<p>生活里最容易理解的是会议室预约。</p>\n<p>同一间会议室，同一段时间，只能被一个团队预约。解决办法不一定是手写锁。可以给数据库加唯一约束，比如会议室 ID 加时间段不能重复。谁先写进去谁成功，后写的人失败。</p>\n<p>这就像登记表上不允许出现两条相同编号的记录。</p>\n<p>库存扣减也类似。</p>\n<p>最容易出错的写法是：</p>\n<pre><code class=\"language-text\">先查库存\n如果库存大于 0\n再扣库存\n</code></pre>\n<p>单人测试没有问题。并发下，两个请求都查到 <code>stock = 1</code>，都觉得可以买，最后就超卖。</p>\n<p>更直接的做法，是把判断和修改合成一条 SQL：</p>\n<pre><code class=\"language-sql\">UPDATE product\nSET stock = stock - 1\nWHERE id = ?\n  AND stock &gt; 0;\n</code></pre>\n<p>影响行数为 1，扣减成功。</p>\n<p>影响行数为 0，库存不足。</p>\n<p>这条 SQL 很朴素，但有用。它把“库存必须大于 0”这个约束放在数据库里执行，避免了“先查再改”中间被别人插队。</p>\n<p>余额更敏感。</p>\n<p>库存超卖，难看，但还能补货、退款、道歉。余额扣错，系统信用就掉了一块。账务系统最怕的不是代码丑，是账对不上。</p>\n<p>余额变更通常不能只想着“加锁”。至少还要考虑事务、流水、幂等和唯一索引。</p>\n<p>事务保证余额更新和流水写入一起成功或失败。</p>\n<p>流水让每一笔变更都有账可查。</p>\n<p>幂等保证同一个扣款请求不会因为重试执行两次。</p>\n<p>唯一索引让重复请求号进不了库。</p>\n<p>这里能看出一个重点：并发控制不是只有锁。</p>\n<p>锁只是其中一种办法。</p>\n<h2>重复提交不是单纯的锁问题</h2>\n<p>用户连续点两次按钮，浏览器重发，客户端超时重试，消息队列重复投递，这些都会带来重复提交。</p>\n<p>比如支付。</p>\n<p>一个用户点了支付，网络卡了一下，他又点一次。或者前端只发了一次，但网关重试了。系统如果每来一次请求都新建一笔支付单，后面就热闹了。</p>\n<p>很多人第一反应是加锁。</p>\n<p>有时可以，但不总是最合适。</p>\n<p>重复提交的关键是：同一件事只能被处理一次。这个问题更接近幂等。</p>\n<p>幂等这个词听起来有点数学，业务里可以简单理解成：同一张凭证来几次，只认第一次，后面返回同一个结果。</p>\n<p>常见做法包括：</p>\n<ul>\n<li>前端按钮置灰，用来改善体验，但不能当后端保障。</li>\n<li>后端使用 <code>idempotency_key</code> 或请求号。</li>\n<li>数据库建立唯一索引，挡住重复请求号。</li>\n<li>重复请求返回第一次处理结果。</li>\n<li>短时间内的重复点击，可以用 Redis 短期锁兜一下。</li>\n</ul>\n<p>锁防的是“同时进来”。</p>\n<p>幂等防的是“同一件事被做多次”。</p>\n<p>两者相关，但不是一回事。</p>\n<h2>任务重复执行：门不在同一间屋子里</h2>\n<p>本地开发时，一个定时任务通常只有一个进程在跑。</p>\n<p>线上部署以后，情况就变了。</p>\n<p>三个实例，每个实例都有同一个 cron。到点以后，三个实例一起扫数据库。原本想处理一批数据，最后处理了三遍。</p>\n<p>这时 Java 的 <code>synchronized</code> 没用。</p>\n<p>因为每个实例都有自己的 JVM、自己的内存、自己的锁。A 机器锁住了，B 机器并不知道。</p>\n<p>单机锁像一间屋子的门锁。多实例部署以后，是好几间屋子里各自挂了一把锁。你锁了自己的门，隔壁屋的人照样在改同一本远程账本。</p>\n<p>这类场景要用大家都看得见的公共门禁。</p>\n<p>常见做法有几种：</p>\n<ul>\n<li>Redis 分布式锁，让同一时间只有一个实例执行。</li>\n<li>数据库任务表抢占，谁更新状态成功谁处理。</li>\n<li>消息队列，把任务拆成消息，由消费者按规则处理。</li>\n<li>调度平台保证同一任务单实例执行。</li>\n</ul>\n<p>这里的重点不是背哪个框架，而是知道“本机锁”和“分布式协调”不是一回事。</p>\n<h2>分布式锁不是豪华版 <code>synchronized</code></h2>\n<p>一说多实例，很多人会自然想到 Redis 分布式锁。</p>\n<p>它确实常用，但别把它想成 <code>synchronized</code> 的远程豪华版。</p>\n<p>分布式锁要考虑很多现实问题：</p>\n<ul>\n<li>锁多久过期？</li>\n<li>业务执行时间超过锁过期时间怎么办？</li>\n<li>服务宕机后锁能不能释放？</li>\n<li>要不要续期？</li>\n<li>释放锁时会不会删掉别人后来拿到的锁？</li>\n</ul>\n<p>最基本的要求是：锁要有过期时间，锁的 value 要有唯一标识，释放时先确认这把锁确实是自己持有的。</p>\n<p>否则可能出现很荒唐的情况。</p>\n<p>A 拿到锁，业务执行太慢，锁过期了。B 拿到新锁，开始处理。A 终于执行完，一把删掉 B 的锁。</p>\n<p>这不是玄学，是真实系统里会发生的普通尴尬。</p>\n<p>所以很多场景下，数据库唯一索引、条件更新、任务表抢占，反而比手写分布式锁更稳。</p>\n<p>分布式锁能用，但要知道自己拿的是什么。</p>\n<h2>状态流转：先把路画出来</h2>\n<p>订单从“待支付”到“已支付”，再到“已发货”，也可能退款、取消、关闭。</p>\n<p>这些状态不是随便改的。</p>\n<p>待支付可以变成已支付，也可以变成已取消。已发货以后就不能假装回到待支付。退款也不是任意时刻都能发生。</p>\n<p>并发情况下，支付回调、用户取消、后台发货、退款请求可能同时到。最后状态乱了，查日志像考古。</p>\n<p>这类问题，锁只能解决一部分。</p>\n<p>更根本的是状态约束。</p>\n<p>常用写法是条件更新：</p>\n<pre><code class=\"language-sql\">UPDATE orders\nSET status = 'PAID'\nWHERE id = ?\n  AND status = 'PENDING_PAYMENT';\n</code></pre>\n<p>只有订单还处于待支付时，才能改成已支付。影响行数为 0，就说明状态已经变了，当前操作不能继续假装成功。</p>\n<p>这就是状态机的朴素版本。</p>\n<p>很多坏代码，就是太自信地 <code>set status = xxx</code>，不管旧状态是什么，直接覆盖。</p>\n<p>状态流转里，锁是一道门，状态机是一张路网。只装门，不画路，还是会迷路。</p>\n<h2>不同语言只是门锁样式不同</h2>\n<p>Java 有 <code>synchronized</code>、<code>ReentrantLock</code>、<code>Atomic</code>。</p>\n<p>Go 有 <code>sync.Mutex</code>、<code>sync.RWMutex</code>，也常用 channel 把共享修改变成排队处理。</p>\n<p>Python 有 <code>threading.Lock</code>，异步代码里有 <code>asyncio.Lock</code>。</p>\n<p>Node.js 主线程是事件循环模型，但服务端仍然会遇到多请求并发、多进程、多实例和数据库并发。</p>\n<p>不要被“Node 是单线程”这句话骗了。</p>\n<p>单线程只能说明一段 JavaScript 代码不会在同一个线程里被两个 CPU 同时执行。它不能阻止两个 HTTP 请求同时更新同一行数据库，也不能阻止两台机器同时跑同一个任务。</p>\n<p>语言层面的锁，更多解决进程内的问题。</p>\n<p>业务系统里的并发，常常发生在数据库、缓存、消息队列、多个服务实例之间。</p>\n<p>所以第一问题不是用什么语言。</p>\n<p>第一问题是：共享数据在哪里？并发入口在哪里？正确性由谁保证？</p>\n<h2>回到 AI</h2>\n<p>AI 会越来越会写代码，也会越来越会规划方案。</p>\n<p>这当然是好事。</p>\n<p>但开发者不能因此把基本概念都交出去。因为系统最后坏不坏，往往不取决于代码长得像不像，而取决于约束有没有放对地方。</p>\n<p>AI 给一个接口加了 <code>synchronized</code>，如果服务是多实例部署，就要知道这把锁可能只管住当前 JVM。</p>\n<p>AI 给库存扣减写了“先查再改”，就要知道这里可能被并发插队。</p>\n<p>AI 给重复提交加了 Redis 锁，也要继续看有没有幂等 key 和唯一索引。</p>\n<p>AI 给订单状态直接覆盖成 <code>PAID</code>，就要问一句：旧状态是什么？这个流转合法吗？</p>\n<p>这不是要每个人都去背 JVM、AQS、JMM、volatile 的细节。那些东西重要，但不是每篇业务代码都要钻到地心。</p>\n<p>更常用、更应该先掌握的是这些判断：</p>\n<ul>\n<li>有没有共享资源？</li>\n<li>有没有并发修改？</li>\n<li>错了以后代价多大？</li>\n<li>能不能失败重试？</li>\n<li>是单进程还是多实例？</li>\n<li>约束放在代码、数据库、Redis、队列，还是调度系统里？</li>\n</ul>\n<p>这些问题想清楚，就能看懂 AI 给出的方案是不是靠谱。</p>\n<p>看不懂这些，AI 写得越快，错误也可能来得越快。</p>\n<p>这话不吓人。它只是工程里的常识：车快了，路标和刹车也要跟上。</p>\n<h2>最后</h2>\n<p>锁解决的是并发修改共享资源时的数据正确性问题。</p>\n<p>悲观锁适合冲突高、不能轻易失败的场景。先关门，再办事。</p>\n<p>乐观锁适合冲突低、允许重试的场景。先办事，提交时验票。</p>\n<p>分布式场景下，本机锁通常不够。数据库锁、Redis 分布式锁、唯一索引、条件更新、消息队列、任务表抢占，都可能是答案的一部分。</p>\n<p>实际业务里，不要一上来问“要不要加锁”。</p>\n<p>先问：</p>\n<ul>\n<li>这段业务并发下会不会出错？</li>\n<li>出错以后影响有多大？</li>\n<li>哪个地方最适合放约束？</li>\n<li>AI 或框架给出的方案，锁住的到底是哪份资源？</li>\n</ul>\n<p>锁不是高级概念。</p>\n<p>它只是提醒人：系统里有些门，不能让所有人一起挤进去。</p>\n<p>至于哪扇门要关，关多久，用钥匙、门禁还是登记簿，那还是人要想明白。</p>\n","date_published":"2026-07-10T00:00:00.000Z","tags":["并发控制","锁","数据库","后端","架构"],"language":"zh"},{"id":"https://www.lihuanyu.com/posts/ai-book-review-beyond-large-prompt/","url":"https://www.lihuanyu.com/posts/ai-book-review-beyond-large-prompt/","title":"AI 整理全书为什么不能只靠一个大 Prompt","summary":"续章的全书整理功能曾把作品设定、章节、事实和记忆一次性发给模型，两次调用跑了八分多钟。后来我把它改成按需查资料的有限 Agent 循环，没用框架，也没再被单次超长上下文拖死。","content_html":"<p>我第一次拿《西游记》测续章时，它把火焰山后面直接接到了灵山受封。</p>\n<p>祭赛国、金光寺、碧波潭、九头虫，这一串本来能写得很热闹的东西，被压得只剩几句交代。我让 AI 往回补，它又开始重复旧章节，像写到半路才想起少交了一段作业。</p>\n<p>续章是我做的 AI 小说应用，英文名是 Novevia，仓库和早期代码里还保留着 Chapterly 这个名字。它会先生成作品设定、故事路线和章节列表，再一章一章往下写。前几章好看不算太难，难的是写到几十章后，它还记得这本书要去哪里。</p>\n<p>当时我以为问题在模型，后来发现，更大的问题是我给模型安排了一种很笨的工作方式。</p>\n<p><img src=\"/assets/posts/2026/ai-book-review-beyond-large-prompt/01-overloaded-book-context.jpg\" alt=\"写作者被整本书的章节卡片和资料淹没，故事路线从火山骤然跳到远方宫殿\"></p>\n<p><a href=\"/en/posts/2026/i-stopped-sending-the-whole-book-to-the-model/\">English version: I Stopped Sending the Whole Book to the Model</a></p>\n<p><a href=\"/posts/from-one-model-call-to-an-agent/\">概念篇：从一次模型调用到一个 Agent</a></p>\n<h2>“全书”两个字把我带偏了</h2>\n<p>最早的实现很符合直觉：既然叫“整理全书”，就把和全书有关的材料都给它。</p>\n<p>作品设定、故事路线、全部章节卡片、事实库、近期记忆和用户要求，服务端先尽量压缩，再拼成一个大 prompt。模型做完诊断，如果服务端检查出危险修改，还要把这批材料连同草案再发一遍，让它自查。</p>\n<p>短作品上，这办法能用。书一长，它的问题就不再是“prompt 写得好不好”，而是一个请求到底能有多胖。</p>\n<p>一次线上任务里，两次模型调用加起来超过了 8 分钟。前端把它改成异步任务，可以避开 HTTP 连接超时，却不能让这 8 分钟消失。页面上的黄条一直在转，我也不知道它是在认真看书，还是已经读晕了。</p>\n<p>更别扭的是，用户有时只问“第 13 章这个标题为什么接不上”，系统仍然把整本书的材料都搬过去。模型看完后，往往给出一句没法反驳、也没什么用的结论：整体节奏偏快，建议加强人物成长。</p>\n<p>材料并不少，缺的是选择。该看哪几章，要不要查故事路线，事实库和记忆对这个问题有没有用，这些判断都被我省掉了。我先把仓库门打开，再叫 AI 自己进去翻，至于要查什么，系统并没有说清楚。</p>\n<h2>先给一张地图</h2>\n<p>后来我把第一步改了。AI 不再先看全部材料，只看一张 manifest，也就是作品清单。</p>\n<p>这张清单里有书名、题材、章节数、已写和未写章节的数量，也会告诉模型故事路线是空的、过薄还是可用，事实库和记忆各有多少条。它不带正文，也不展开整份大纲。</p>\n<p><img src=\"/assets/posts/2026/ai-book-review-beyond-large-prompt/02-manifest-first-retrieval.jpg\" alt=\"从庞大的资料库中先看索引，再按问题取出少量相关档案\"></p>\n<p>模型看完清单，才选择工具。它可以查某一段章节卡片，读路线摘要，取某一章的摘要、开头或结尾，也可以按关键词查事实库和章节记忆。最近的调整记录也做成了工具，免得它隔几天又提一遍同样的建议。</p>\n<p>我没有用 provider 自带的 function calling，只约定了一个 JSON 格式：</p>\n<pre><code class=\"language-json\">{\n  &quot;type&quot;: &quot;tool_call&quot;,\n  &quot;tool&quot;: &quot;list_chapter_cards&quot;,\n  &quot;args&quot;: { &quot;fromChapterNo&quot;: 9, &quot;toChapterNo&quot;: 15 },\n  &quot;reason&quot;: &quot;需要确认祭赛国前后的章节承接&quot;\n}\n</code></pre>\n<p>服务端收到 <code>tool_call</code> 后执行工具，截断过长字段，再把结果放回对话。模型可以继续查，也可以认为证据已经足够，输出 <code>final</code> 方案。</p>\n<p>这样一来，“第 13 章标题不对”只需查相邻几章和路线；“祭赛国写得太快”才需要进一步搜索金光寺、碧波潭和九头虫。整本书的长度不再直接决定第一个请求的长度，问题本身开始决定该读多少材料。</p>\n<h2>这个循环算不算 ReAct</h2>\n<p>代码写完后，我才回头想这个问题。</p>\n<p>ReAct 是 Reasoning and Acting 的缩写。模型不做一次性回答，它在推理、行动和观察之间往返：决定查什么，调用工具，看到结果，再决定下一步。</p>\n<p>续章的实现符合这个核心。模型根据上一次 observation 选择下一个工具，代码没有预先写死顺序。它也没有复制论文里的完整 Thought 格式，不保存模型的长思维链，只记一句简短的 <code>reason</code>。所以我更愿意把它叫做有限的 ReAct-style loop。</p>\n<p>不过，ReAct 并不是这轮改造里最重要的那个词。它只是一个后来对得上的名字。对产品更有用的变化，是上下文从“服务端事先打包好的全部材料”，变成了“模型围绕当前问题建立的工作集”。</p>\n<p>在这个功能里，这也是 Agent 和普通模型调用的分界线。模型是其中的决策部件，Agent 还得有工具、状态、循环、停止条件和权限边界。少了后面这些东西，只是一个会输出 JSON 的模型。</p>\n<h2>Agent 不能直接改书</h2>\n<p>一旦让模型自己选工具，下一个问题就是：它可以走多远。</p>\n<p>我给的答案很保守。不同模型档位只允许 3 到 5 个决策步骤，工具也只能调用 3 到 5 次。每次请求有 24,000 到 64,000 字符的上限，超出后保留系统指令、manifest 和最近的工具结果。这不是精确的 token 预算，但能防止对话随着工具调用一路变胖。</p>\n<p><img src=\"/assets/posts/2026/ai-book-review-beyond-large-prompt/03-bounded-agent-loop.jpg\" alt=\"有限的资料查询循环经过步数限制、只读档案、方案校验和人工确认\"></p>\n<p>更重要的限制是，所有查资料的工具都只读。模型不能写数据库，只能提交一份结构化方案。<code>validate_draft_plan</code> 会检查它是不是在覆盖已写正文，是不是把新章节插到了不安全的位置，或者一口气补了过多章节。检查不通过，模型最多修正一次。</p>\n<p>已经有正文的章节只会收到建议，未写且为 0 字的计划章才能进入自动调整列表。用户看完方案后还要点一次确认，应用后也可以回滚。让模型直接改库确实能少写几段代码，只是这几段代码迟早会从事故里补回来。</p>\n<p>运行方式也从一次 HTTP 请求变成了持久化任务。页面可以关，任务可以取消；后台进程中断后，worker 会把超时的 <code>running</code> 任务重新排队，最多恢复两次。每个 <code>tool_called</code>、<code>tool_result</code> 和 <code>plan_validated</code> 都会记到任务事件里，前端会显示最近一步，刷新后也能继续看。</p>\n<h2>Pi 最后没有接</h2>\n<p>早期方案里，我考虑过接 <code>@earendil-works/pi-ai</code>。等到真正开始写全书整理，这个依赖并没有进项目。</p>\n<p><code>pi-ai</code> 更靠近模型接入层：统一 provider、消息格式、流式输出、tool call 和 usage。如果连 Agent runtime 也用 Pi 的实现，还能少写一部分消息循环、状态和工具执行代码。但章节怎么查，事实库怎么搜，哪些内容绝对不能自动改，这些仍然是续章自己的事。</p>\n<p>项目当时已经有 <code>NovelAiGatewayService</code>，provider 选择、错误归一、usage 和成本记录都在这里。我没有为了再做一遍这些事而换掉它，只在上面加了 <code>NovelBookDoctorRuntime</code>，再用 <code>NovelBookDoctorToolRegistry</code> 收拢工具。任务的排队、恢复、应用和回滚，继续放在 <code>NovelBookAdjustmentService</code> 里。</p>\n<p>所以，这个 Agent 的编排层算是手写的，底下的模型调用和任务基础设施不是。我不反对 Agent 框架，只是这个循环最多 5 步，工具数量也不多，还没有复杂到值得替换既有调用链。等多个 Agent 开始共用工具、记忆、trace 和重试策略时，账可能会另算。</p>\n<h2>超时没了，账还没算完</h2>\n<p>改完以后，我最想解决的问题已经解决了：整理全书没有再因为单次提交超长上下文而超时。</p>\n<p>超时仍然可能发生。provider 的响应会变慢，网络会出问题，整个任务也有时限。现在消失的是一个明确的故障模式：书越长，第一个请求越长，最后把自己拖死。</p>\n<p>总 token 也不一定比原来少。ReAct 会把一次大调用拆成几次小调用，如果模型连续查了路线、章节、事实和记忆，加起来可能更贵。现在能确定的是单次请求受控，不能直接推出“token 大幅下降”。</p>\n<p>质量也一样。目前的测试能证明 manifest 不带完整内容，章节卡片不泄漏正文，单章工具只返回请求的片段，危险草案会被拦下。这些测试能保住边界，不能证明 AI 已经有了资深编辑的判断力。真要比质量，还得固定一批作品和问题，记录方案采纳率、误修率和实际工具路径。</p>\n<p>现在再点一次“整理全书”，页面上的黄条还是会转。不同的是，任务记录里能看到它刚查了哪几章，为什么又去翻事实库，方案如果被拦，又是卡在哪条规则上。我不用再盯着一个八分钟没动静的请求，猜它究竟在忙，还是已经被我塞晕了。</p>\n<p>至于 ReAct，这个名字放在代码说明里很合适，放在这轮实践里倒没那么神秘。模型没有突然读懂《西游记》，我也没找到一个无所不能的 Agent 框架。它只是照着目录，缺什么拿什么；拿够了，停下来交方案。</p>\n<p>黄条照样会转，好在现在能转到头了。</p>\n","date_published":"2026-06-21T00:00:00.000Z","date_modified":"2026-08-08T00:00:00.000Z","tags":["AI","Agent","ReAct","Prompt","上下文工程"],"language":"zh"},{"id":"https://www.lihuanyu.com/en/posts/2026/i-stopped-sending-the-whole-book-to-the-model/","url":"https://www.lihuanyu.com/en/posts/2026/i-stopped-sending-the-whole-book-to-the-model/","title":"I Stopped Sending the Whole Book to the Model","summary":"My whole-book review feature used to send the setting, chapter plan, facts, and memories in one giant prompt. Two model calls took more than eight minutes. I replaced that request with a small, bounded agent loop that retrieves context only when it needs it.","content_html":"<p>The first time I tested Novevia with <em>Journey to the West</em>, it jumped straight from the Flaming Mountains to the final arrival at Spirit Mountain.</p>\n<p>The Kingdom of Jisai, Golden Light Monastery, Emerald Wave Lake, the Nine-Headed Beast: a run of adventures that should have filled several lively chapters was reduced to a few lines. When I asked the AI to fill the gap, it began repeating earlier chapters. It felt like a student who remembered, halfway through an assignment, that several pages were missing.</p>\n<p>Novevia is an AI fiction-writing app I built. Some early code still uses its old name, Chapterly. It creates the setting, story route, and chapter plan before writing the book one chapter at a time. Making the first few chapters interesting is not the hardest part. The hard part is reaching chapter fifty without forgetting where the book is going.</p>\n<p>At first I blamed the model. The larger problem turned out to be the job I had given it.</p>\n<p><img src=\"/assets/posts/2026/ai-book-review-beyond-large-prompt/01-overloaded-book-context.jpg\" alt=\"A writer buried under chapter cards and reference material while the story route jumps from a volcano to a distant palace\"></p>\n<p><a href=\"/posts/ai-book-review-beyond-large-prompt/\">Chinese version of this article</a></p>\n<p><a href=\"/en/posts/2026/from-one-model-call-to-an-agent/\">Concept guide: From One Model Call to an Agent</a></p>\n<h2>“Whole Book” Sent Me in the Wrong Direction</h2>\n<p>The original implementation sounded reasonable. If the feature was called “whole-book review,” I should give it the whole book.</p>\n<p>The server collected the setting, story route, every chapter card, the fact store, recent memories, and the user’s instructions. It compressed what it could, joined everything into one giant prompt, and sent the lot to the model. If a server-side check found a risky change in the draft plan, the second call included most of that material again so the model could revise its answer.</p>\n<p>This worked for short books. As a book grew, prompt wording stopped being the main problem. The request itself had become too large.</p>\n<p>In one production task, two model calls took more than eight minutes in total. Moving the operation into an asynchronous job prevented the HTTP connection from timing out, but it did not make those eight minutes disappear. A yellow progress strip kept spinning in the UI. I could not tell whether the model was carefully reading the book or had gone cross-eyed somewhere in the prompt.</p>\n<p>The waste was hard to ignore. A user might ask, “Why does chapter 13 no longer connect to chapter 12?” The system would still carry the entire archive to the model. The model could read all of it and return a conclusion that was difficult to dispute and almost impossible to use: the overall pace is too fast; consider strengthening character growth.</p>\n<p>There was plenty of context. What was missing was selection.</p>\n<p>Which chapters matter? Does this question require the story route? Are facts and memories relevant? I had removed those decisions from the process. I opened the warehouse door and told the AI to look around without first deciding what it was looking for.</p>\n<h2>Start With an Index</h2>\n<p>I changed the first step. The model no longer receives the full set of material. It receives a manifest.</p>\n<p>The manifest contains the title, genre, chapter count, and the number of written and unwritten chapters. It says whether the story route is empty, thin, or usable, and how many facts and memories are available. It contains no prose and does not expand the full outline.</p>\n<p><img src=\"/assets/posts/2026/ai-book-review-beyond-large-prompt/02-manifest-first-retrieval.jpg\" alt=\"A compact index guides the retrieval of a few relevant folders from a much larger archive\"></p>\n<p>After reading that index, the model chooses a tool. It can inspect a range of chapter cards, read a route summary, fetch the summary or opening or ending of one chapter, and search the fact store or chapter memories by keyword. Recent adjustment records are another tool, so the model does not return a few days later with the same advice.</p>\n<p>I did not use the provider’s native function calling. The runtime uses a small JSON protocol instead:</p>\n<pre><code class=\"language-json\">{\n  &quot;type&quot;: &quot;tool_call&quot;,\n  &quot;tool&quot;: &quot;list_chapter_cards&quot;,\n  &quot;args&quot;: { &quot;fromChapterNo&quot;: 9, &quot;toChapterNo&quot;: 15 },\n  &quot;reason&quot;: &quot;Check how the chapters around the Kingdom of Jisai connect&quot;\n}\n</code></pre>\n<p>The server executes the tool, truncates fields that are too long, and returns the observation to the conversation. The model can query again or decide that it has enough evidence and produce a <code>final</code> plan.</p>\n<p>Now a question about chapter 13 may require only its neighboring chapter cards and the route. A complaint that the Kingdom of Jisai was rushed may justify another search for Golden Light Monastery, Emerald Wave Lake, and the Nine-Headed Beast. The length of the book no longer determines the size of the first request. The question determines how much material gets opened.</p>\n<h2>Does This Count as ReAct?</h2>\n<p>I only asked that question after the code was working.</p>\n<p>ReAct stands for Reasoning and Acting. Instead of producing a one-shot answer, the model moves between deciding, acting, and observing: choose what to inspect, call a tool, read the result, then choose the next step.</p>\n<p>Novevia’s implementation fits that core idea. The model selects each tool after seeing the previous observation; the order is not hard-coded by the server. It does not copy the full Thought format from the paper or store a long chain of thought. The tool call keeps only a short <code>reason</code>. I think “bounded ReAct-style loop” is an accurate description.</p>\n<p>ReAct, however, was not the important part of the change. It was a name I found after the fact. The useful change was turning context from a package assembled in advance by the server into a working set built around the current question.</p>\n<p>That is also where this feature becomes more than an ordinary model call. The model is the decision-making component. The agent also needs tools, state, a loop, stopping conditions, and permission boundaries. Without those pieces, it is only a model that knows how to emit JSON.</p>\n<h2>The Agent Does Not Get to Edit the Book</h2>\n<p>Once the model can choose its own tools, the next question is how far it may go.</p>\n<p>My answer is conservative. Depending on the model tier, the runtime allows only three to five decision steps and three to five tool calls. Each request has a prompt budget between 24,000 and 64,000 characters. If the conversation exceeds that budget, the runtime keeps the system instructions, manifest, and most recent tool results. Character count is not a precise token budget, but it stops the conversation from growing without bound.</p>\n<p><img src=\"/assets/posts/2026/ai-book-review-beyond-large-prompt/03-bounded-agent-loop.jpg\" alt=\"A bounded retrieval loop passes through step limits, read-only archives, validation, and human approval\"></p>\n<p>More importantly, every retrieval tool is read-only. The model cannot write to the database. It can only submit a structured plan.</p>\n<p><code>validate_draft_plan</code> checks whether that plan would overwrite written prose, insert chapters in an unsafe place, or add too many chapters at once. If validation fails, the model gets at most one revision.</p>\n<p>Chapters that already contain prose receive suggestions only. Only planned chapters with zero words can enter the automatic adjustment list. The user still reviews the plan and confirms it before anything is applied, and an applied change can be rolled back. Letting the model write directly would have saved a little code. That code would eventually have been written in response to an incident instead.</p>\n<p>The operation also moved from a single HTTP request into a persistent task. Closing the page no longer kills the task, and the UI exposes a cancel action. After a worker interruption, the service requeues stale <code>running</code> tasks and allows up to two recovery attempts. The worker records events such as <code>tool_called</code>, <code>tool_result</code>, and <code>plan_validated</code>, so the UI can show what happened even after a refresh.</p>\n<h2>Why I Did Not Adopt Pi</h2>\n<p>An early design considered <code>@earendil-works/pi-ai</code>. It never became a project dependency.</p>\n<p><code>pi-ai</code> sits closer to the model integration layer. It can normalize providers, message formats, streaming, tool calls, and usage. Using Pi’s agent runtime could also remove some hand-written message-loop and tool-execution code. It would not decide how chapter lookup should work, how the fact store should be searched, or which parts of a book must never be changed automatically. Those rules still belong to Novevia.</p>\n<p>The project already had <code>NovelAiGatewayService</code> for provider selection, normalized errors, usage, and cost records. Replacing that layer would have duplicated work. I added <code>NovelBookDoctorRuntime</code> above it and collected the read tools in <code>NovelBookDoctorToolRegistry</code>. Queueing, recovery, application, and rollback stayed in <code>NovelBookAdjustmentService</code>.</p>\n<p>So the orchestration layer is hand-written, while the model gateway and task infrastructure are not. I have nothing against agent frameworks. This loop has at most five steps and only a small set of tools; replacing the existing call path would have cost more than it removed. If several agents later need to share tools, memory, traces, and retry policy, that calculation may change.</p>\n<h2>The Timeout Is Gone. The Accounting Is Not Finished.</h2>\n<p>The failure I most wanted to remove is gone. Whole-book review no longer times out because the first request contains an ever-growing copy of the book.</p>\n<p>Other timeouts remain possible. A provider can slow down, the network can fail, and the task itself has a deadline. What disappeared was one specific failure mode: the longer the book became, the larger the first request became, until it sank under its own context.</p>\n<p>Total token usage may not be lower. ReAct turns one large call into several smaller calls. If the model inspects the route, chapters, facts, and memories in succession, the total may even cost more. I can claim that each request is bounded. I cannot turn that fact into “token usage dropped dramatically.”</p>\n<p>The same caution applies to quality. Tests prove that the manifest does not contain full content, chapter-card tools do not leak prose, a chapter tool returns only the requested fragment, and dangerous plans are rejected. Those tests protect boundaries. They do not prove that the AI has acquired the judgment of an experienced editor. Measuring that will require a fixed set of books and questions, followed by plan acceptance rates, incorrect-edit rates, and the actual tool paths taken.</p>\n<p>When I start a whole-book review now, the yellow strip still spins. The difference is that the task log shows which chapters the agent inspected, why it opened the fact store, and which rule stopped a plan. I no longer stare at an eight-minute request and wonder whether it is busy or merely buried.</p>\n<p>ReAct is a useful name in the code documentation. In this project, it is less mysterious than it sounds. The model did not suddenly understand <em>Journey to the West</em>, and I did not discover an all-powerful agent framework. It follows an index, fetches what it lacks, and stops when it has enough to submit a plan.</p>\n<p>The yellow strip still spins. At least now it reaches the end.</p>\n","date_published":"2026-06-21T00:00:00.000Z","date_modified":"2026-08-08T00:00:00.000Z","tags":["AI","Agent","ReAct","Prompt","Context Engineering"],"language":"en"},{"id":"https://www.lihuanyu.com/en/posts/2026/oidc-identity-boundary-behind-login/","url":"https://www.lihuanyu.com/en/posts/2026/oidc-identity-boundary-behind-login/","title":"OIDC: The Identity Boundary Behind Login","summary":"A practical introduction to OpenID Connect: what problem it solves, when it is worth using, how it fits login and payments, and how it differs from OAuth, SAML, CAS, LDAP, JWT, and passkeys.","content_html":"<p>I recently went back and cleaned up the login systems around a few small applications. The more I worked through it, the more I felt that most applications do not have that many truly unavoidable problems.</p>\n<p>The first one is login: who the user is, how they enter the system, and what identity they have after they enter.</p>\n<p>The second one is payment: why the user pays, how much they paid, and what they should receive in the product after paying.</p>\n<p>Other features are important, of course. They may decide whether a product has any value at all. But if login and payment are blurry, an application is like a small roadside shop with nobody watching the door and nobody keeping the ledger. It may look busy during the day. The trouble appears when it is time to close the books at night.</p>\n<p>In the AI era, these two problems become even more concrete. In the old web, a free tool being clicked a few more times usually meant some bandwidth and server load. In many AI applications, a real use may trigger model calls, credit consumption, and a visible bill. Can anonymous users use it? Who owns the free quota? Where do paid plans and permissions live? If a user comes from another application, are they the same person?</p>\n<p>These questions eventually return to a plain word: identity.</p>\n<p>OIDC is one standard way to answer it.</p>\n<p><img src=\"/assets/posts/2026/oidc-login-identity/01-identity-station.jpg\" alt=\"A quiet identity station issuing passes for several applications\"></p>\n<p><a href=\"/posts/oidc-identity-boundary-behind-login/\">Chinese version of this article</a></p>\n<h2>What OIDC Is</h2>\n<p>OIDC stands for OpenID Connect. It is built on OAuth 2.0 and is used for authentication.</p>\n<p>In one sentence:</p>\n<pre><code class=\"language-text\">OAuth 2.0 mainly answers: can this client access a resource?\nOIDC mainly answers: who is the current logged-in user?\n</code></pre>\n<p>When people say “Log in with Google”, “Log in with GitHub”, or “use the company SSO account”, they are often not talking about OAuth alone. They are talking about identity on top of OAuth. That identity layer is usually OIDC, or at least something very close to it.</p>\n<p>The most important output of OIDC is the <code>id_token</code>. It is usually a JWT containing claims like these:</p>\n<pre><code class=\"language-json\">{\n  &quot;iss&quot;: &quot;https://accounts.example.com&quot;,\n  &quot;sub&quot;: &quot;acct_123&quot;,\n  &quot;aud&quot;: &quot;my-web-app&quot;,\n  &quot;exp&quot;: 1780000000,\n  &quot;email&quot;: &quot;user@example.com&quot;,\n  &quot;email_verified&quot;: true\n}\n</code></pre>\n<p>Three fields matter especially:</p>\n<ul>\n<li><code>iss</code>: who issued this identity result.</li>\n<li><code>sub</code>: the stable user identifier at the issuer.</li>\n<li><code>aud</code>: which application this identity result is intended for.</li>\n</ul>\n<p>After an application receives an <code>id_token</code>, it should not glance at the email address and call the user logged in. It has to verify the signature, verify <code>iss</code>, verify <code>aud</code>, verify the expiration time, and then use <code>iss + sub</code> to find or create its own local user.</p>\n<p>That sounds fussy, but it solves a fundamental problem. Identity should not be whatever the frontend says. It should not come from a URL parameter. It should not be guessed independently by each application. Identity is issued by a known authority, and applications verify it by a standard process.</p>\n<p>That is the boundary.</p>\n<h2>Why Not Just Write a Login API</h2>\n<p>A small application often starts with something like this:</p>\n<pre><code class=\"language-text\">POST /login\n  email + code\n  -&gt; returns token\n</code></pre>\n<p>This can work. I do not think every early project needs a grand identity platform before it has users. Many projects do not die because they are insufficiently standard. They die because they build a cathedral before opening the front door.</p>\n<p>The problem appears when there is more than one application.</p>\n<p>The first application wants email login. The second wants GitHub login. The third wants Google. The fourth needs a mini program login. The admin console needs login too. The CLI also needs login. A separate product wants to reuse the same account.</p>\n<p>If every application builds this again by itself, a few things happen quickly:</p>\n<ul>\n<li>A user logs in to application A, then has to register again in application B.</li>\n<li>The same email becomes different users in different products.</li>\n<li>Third-party callbacks, secrets, and state checks are scattered everywhere.</li>\n<li>Email codes, rate limits, risk control, bans, and audit logs are rebuilt several times.</li>\n<li>Later, when unified membership, cross-application entitlements, or account merging becomes necessary, the foundation is already loose.</li>\n</ul>\n<p>The login endpoint itself is not complicated. Maintaining identity relationships over time is.</p>\n<p>This is where OIDC earns its keep. It does not make the login button prettier. It separates the question of “who is allowed to prove the user’s identity” from the business logic of each application.</p>\n<p>A central identity service handles authentication: who the user is, which method they used, whether the email is verified, whether the account is disabled. Each application handles its own business: what this user is called here, what role they have, what plan they bought, how many credits remain, whether they are overdue.</p>\n<p>Those two concerns should not be kneaded into one lump.</p>\n<h2>Central Identity and Local Users</h2>\n<p>When building unified login, it is easy to run to the other extreme: if there is a central identity service, should all applications share one user table?</p>\n<p>My answer is no.</p>\n<p>A steadier structure looks like this:</p>\n<pre><code class=\"language-text\">Identity center\n  identity: acct_123\n  email: user@example.com\n  providers: email / google / github\n\nApplication A\n  user_id: 1\n  auth_issuer: https://accounts.example.com\n  auth_subject: acct_123\n  role: admin\n  plan: pro\n\nApplication B\n  user_id: 58\n  auth_issuer: https://accounts.example.com\n  auth_subject: acct_123\n  credits: 1200\n  status: active\n</code></pre>\n<p>The identity center answers: this is the same person.</p>\n<p>The local user table answers: this is the person’s state inside this application.</p>\n<p>That division matters.</p>\n<p>Login is usually a global problem. Payment and permissions are often application-specific problems. If someone is a paid member in a writing tool, that does not mean they should have the same quota in an image tool. If someone is an administrator in one console, that does not mean they should be an administrator somewhere else. If the central account is disabled, every application should stop access. But a business-level ban inside one application may not need to affect the user’s use of another product.</p>\n<p>If everything is stuffed into the <code>id_token</code>, the token eventually becomes a small database. It looks powerful. It is usually dangerous.</p>\n<p><code>id_token</code> is good for identity facts. It is not a good home for fast-changing business state. Plans, credits, points, orders, roles, ban reasons, and usage limits should normally be read from the application’s own system.</p>\n<p><img src=\"/assets/posts/2026/oidc-login-identity/03-local-users-and-billing.jpg\" alt=\"One central identity mapped to several application-specific user ledgers\"></p>\n<h2>A Minimal Login Flow</h2>\n<p>OIDC has many terms, but the common web flow can first be understood through Authorization Code + PKCE.</p>\n<p>Roughly:</p>\n<pre><code class=\"language-text\">1. The user clicks login in the application.\n2. The application creates state, nonce, code_verifier, and code_challenge.\n3. The application redirects the user to the identity center's /authorize endpoint.\n4. The user logs in at the identity center.\n5. The identity center redirects back with a temporary code.\n6. The application verifies state.\n7. The application exchanges code + code_verifier for tokens at /token.\n8. The application verifies id_token.\n9. The application finds or creates a local user by iss + sub.\n10. The application creates its own session.\n</code></pre>\n<p>Several words are easy to mix up.</p>\n<p><code>code</code> is a one-time temporary ticket. It is not a login session.</p>\n<p><code>id_token</code> is the identity result. It proves who the user is.</p>\n<p><code>access_token</code> is a credential for accessing resources. It should not casually be treated as proof of user identity.</p>\n<p><code>refresh_token</code> is a renewal credential. Whether to issue it, to whom, how long it lives, and how it rotates all need care.</p>\n<p>For an ordinary web application, the comfortable shape is often this: the backend handles the callback and token exchange, verifies <code>id_token</code>, then writes its own HttpOnly cookie for the browser. The browser does not need to carry long-lived identity-provider tokens around in frontend code.</p>\n<p><img src=\"/assets/posts/2026/oidc-login-identity/02-code-token-flow.jpg\" alt=\"A simplified OIDC path where code returns to the backend and tokens are exchanged server-side\"></p>\n<p>A minimal pseudo implementation looks like this:</p>\n<pre><code class=\"language-js\">app.get('/auth/start', (req, res) =&gt; {\n  const state = randomString();\n  const nonce = randomString();\n  const verifier = randomString();\n  const challenge = sha256base64url(verifier);\n\n  saveTemporaryLoginState(req, { state, nonce, verifier });\n\n  res.redirect(\n    'https://accounts.example.com/authorize?' +\n      new URLSearchParams({\n        response_type: 'code',\n        client_id: 'my-web-app',\n        redirect_uri: 'https://app.example.com/auth/callback',\n        scope: 'openid email profile',\n        state,\n        nonce,\n        code_challenge: challenge,\n        code_challenge_method: 'S256',\n      }),\n  );\n});\n\napp.get('/auth/callback', async (req, res) =&gt; {\n  const saved = readTemporaryLoginState(req);\n  if (req.query.state !== saved.state) {\n    return res.status(400).send('bad state');\n  }\n\n  const tokens = await exchangeCodeForTokens({\n    code: req.query.code,\n    codeVerifier: saved.verifier,\n  });\n\n  const claims = await verifyIdToken(tokens.id_token, {\n    issuer: 'https://accounts.example.com',\n    audience: 'my-web-app',\n    nonce: saved.nonce,\n  });\n\n  const user = await findOrCreateLocalUser({\n    issuer: claims.iss,\n    subject: claims.sub,\n    email: claims.email,\n  });\n\n  writeLocalSessionCookie(res, user.id);\n  res.redirect('/dashboard');\n});\n</code></pre>\n<p>Real projects also need error handling, expiration handling, redirect allowlists, secure cookie attributes, logging, rate limiting, and a few unpleasant edge cases. But the main line is this.</p>\n<h2>When OIDC Fits</h2>\n<p>OIDC is useful in several situations.</p>\n<p>First, multiple applications need shared login.</p>\n<p>Even if there are only two applications today, it is worth thinking about this early. Once login spreads out, pulling it back later is expensive. User data, third-party identities, email verification, old tokens, historical orders, and support records all become migration work.</p>\n<p>Second, login and business users should be separated.</p>\n<p>This is especially important for products with payments, credits, quotas, teams, roles, bans, or other application-specific state. The identity center should not become a super business-user table. The local user table should not rebuild a full identity-provider system either.</p>\n<p>Third, third-party identity providers need to be connected.</p>\n<p>Google, Microsoft, enterprise identity, and organization accounts all fit naturally into the OIDC model. Even when a platform document calls the feature “OAuth login”, there is often an <code>id_token</code> or a similar identity-verification step behind it.</p>\n<p>Fourth, CLI, mobile apps, and separate services need a common login path.</p>\n<p>Web applications can use Authorization Code + PKCE. A CLI can also use PKCE if a browser is available. Without a browser, device code flow may be a better fit. Mobile apps should use the system browser or the platform-recommended approach. Asking users to type passwords into an untrusted WebView is asking for trouble.</p>\n<p>Fifth, cross-application membership or account governance may appear later.</p>\n<p>Payment does not always belong inside the identity center. But the identity center should at least answer one question stably: which person is this? Without that, orders, entitlements, credits, risk control, and support cases are hard to coordinate across products.</p>\n<h2>When Not to Rush</h2>\n<p>Not every project needs OIDC immediately.</p>\n<p>If it is a small team tool with one web frontend and a small user base, simple session login may be enough.</p>\n<p>If the application is completely tied to one platform, such as a mini program where the user system revolves around that platform’s openid, it may be more practical to make that platform login solid first.</p>\n<p>If the problem is service-to-service communication, the question is usually not “who is the user” but “is this service allowed to call that service”. API keys, mTLS, or OAuth client credentials may fit better.</p>\n<p>If the goal is to let users avoid passwords, WebAuthn and passkeys are excellent authentication methods, but they do not replace OIDC. A passkey is closer to a way for the user to prove themselves. OIDC is a way for applications to trust the identity result issued by an identity provider.</p>\n<p>There is another case worth watching: adopting a heavy identity platform too early just to look professional. The business has not run yet, but callbacks, certificates, clients, realms, and scopes have already wrapped the team into a knot.</p>\n<p>Standards are meant to reduce long-term complexity. They should not manufacture complexity on day one.</p>\n<h2>How It Differs From Other Options</h2>\n<p>OAuth 2.0 and OIDC are the easiest pair to confuse.</p>\n<p>OAuth is about authorization. For example, an application wants to read a user’s cloud drive files. The user grants permission, then the application uses an access token to call the drive API. The key question is: can this client access this resource?</p>\n<p>OIDC adds an identity layer on top of OAuth. It standardizes <code>id_token</code>, <code>sub</code>, <code>iss</code>, <code>aud</code>, <code>nonce</code>, Discovery, JWKS, and UserInfo. It lets applications verify who the logged-in user is.</p>\n<p>SAML is older and still common, especially in enterprise SSO. It is XML-based and has a mature enterprise software ecosystem. For modern web and mobile applications, though, the developer experience is usually lighter with OIDC.</p>\n<p>CAS appears often in universities and older organization systems. It provides straightforward single sign-on, but its ecosystem is not as common as OIDC for modern API and app scenarios.</p>\n<p>LDAP and Active Directory are closer to directory services. They are good at storing organizations, users, groups, and credentials, and they are widely used in enterprises. But letting every modern web application connect directly to LDAP is usually rougher than putting an OIDC provider in front.</p>\n<p>Session cookies are the most common login state for a single web application. They do not conflict with OIDC. In many cases, the clean design is exactly this: use OIDC for external login, then use an application-owned HttpOnly session cookie for local login state.</p>\n<p>JWT is not an alternative to OIDC. JWT is only a token format. Signing your own JWT does not mean you have implemented OIDC. OIDC cares about issuer, audience, discovery, public keys, flows, and verification rules.</p>\n<p>WebAuthn and passkeys solve “how the user proves themselves”. OIDC solves “how an application trusts an identity result from an identity provider”. They can work together: the user logs in to the identity center with a passkey, and applications receive the result through OIDC.</p>\n<h2>Why These Concepts Matter More in the AI Era</h2>\n<p>AI makes code faster to write. It also makes system design problems show up earlier.</p>\n<p>Before, an application could survive for quite a while with one user table and a few endpoints. Now applications connect more easily to each other. Admin consoles, CLIs, agents, mobile apps, independent services, and automation jobs can all become entrances to the same product capability.</p>\n<p>AI also lets more people build more of the system by themselves. One person can write frontend, backend, admin pages, deployment scripts, payment flows, and login flows. Once that becomes possible, the dangerous part is not failing to write code. The dangerous part is not knowing where the boundaries should be.</p>\n<p>Login is especially like this.</p>\n<p>If login is modeled badly, it is not a matter of changing a few pages later. User identity, payment records, permissions, risk control, account merging, deletion, and auditing all get tied together. AI can help write the code. The concepts still need to be understood by the person building the system.</p>\n<p>For independent developers and small teams, a few low-level judgments are worth keeping:</p>\n<ul>\n<li>Login is an identity boundary, not just a button.</li>\n<li>Payment is a ledger of business entitlements, not just a webhook.</li>\n<li>The identity center should not swallow the application user table.</li>\n<li>The application user table should not rebuild the identity center.</li>\n<li>More tokens are not automatically better. Larger scopes are not automatically better.</li>\n<li>Who the user is and what the user can do should be modeled separately.</li>\n</ul>\n<p>These judgments are not flashy. They are useful.</p>\n<h2>Finally</h2>\n<p>OIDC is not something every project must adopt on the first day.</p>\n<p>But once there are multiple applications, once the same person should have one identity across products, or once repeated email login, third-party callbacks, verification codes, risk control, and account merging start appearing in different places, it is time to understand OIDC seriously.</p>\n<p>It does not solve “how to draw a login page”. It solves “who is allowed to prove who this user is”.</p>\n<p>Once that question is clear, many later designs become calmer: the identity center handles authentication, applications handle business state; <code>id_token</code> proves identity, local sessions keep users logged in; <code>iss + sub</code> gives a stable mapping, while payments and permissions stay in each application’s own ledger.</p>\n<p>Technical protocols eventually return to ordinary bookkeeping.</p>\n<p>Someone has to watch the door. Someone has to keep the accounts. If an application wants to live for a long time, login and payment cannot stay vague.</p>\n","date_published":"2026-06-20T00:00:00.000Z","tags":["OIDC","OAuth","Login","Authentication","Architecture"],"language":"en"},{"id":"https://www.lihuanyu.com/posts/oidc-identity-boundary-behind-login/","url":"https://www.lihuanyu.com/posts/oidc-identity-boundary-behind-login/","title":"OIDC：登录背后的身份边界","summary":"从个人产品和 AI 应用的真实工程需求出发，聊聊 OpenID Connect 解决的到底是什么问题、什么时候适合使用、和 OAuth、SAML、CAS、LDAP、WebAuthn 等方案有什么区别。","content_html":"<p>最近把几个小应用的登录体系重新梳理了一遍，越做越觉得，很多应用真正绕不开的东西其实不多。</p>\n<p>第一是登录。用户是谁，怎么进来，进来以后在系统里是什么身份。</p>\n<p>第二是付费。用户为什么付钱，付了多少钱，付完以后在产品里能得到什么。</p>\n<p>剩下的功能当然重要，甚至决定产品有没有价值。但如果登录和付费没有想清楚，应用就像一间开在路边的小店：门口没人看，账本也没人记。白天也许热闹，晚上盘账时就知道问题来了。</p>\n<p>AI 时代以后，这两个问题反而更重要。以前一个网页工具被人多点几次，最多是带宽和服务器压力。现在很多 AI 应用每一次真正使用，后面都有模型调用、额度消耗和成本账单。未登录用户能不能用？免费额度算给谁？付费用户的套餐和权限放在哪里？一个用户从另一个应用过来，算不算同一个人？</p>\n<p>这些问题最后都会回到一个很朴素的地方：身份。</p>\n<p>OIDC 就是解决这个问题的一套标准方法。</p>\n<p><img src=\"/assets/posts/2026/oidc-login-identity/01-identity-station.jpg\" alt=\"统一身份中心像一个安静的通行证签发处\"></p>\n<p><a href=\"/en/posts/2026/oidc-identity-boundary-behind-login/\">English version: OIDC: The Identity Boundary Behind Login</a></p>\n<h2>先说清楚：OIDC 是什么</h2>\n<p>OIDC，全称 OpenID Connect。它建立在 OAuth 2.0 之上，用来做身份认证。</p>\n<p>一句话说：</p>\n<pre><code class=\"language-text\">OAuth 2.0 主要回答：这个客户端能不能访问某个资源？\nOIDC 主要回答：当前登录的这个人是谁？\n</code></pre>\n<p>平时我们说“使用 Google 登录”“使用 GitHub 登录”“公司统一账号登录”，严格讲，很多时候说的都不只是 OAuth，而是 OAuth 之上的身份协议，也就是 OIDC。</p>\n<p>OIDC 最关键的产物是 <code>id_token</code>。它通常是一个 JWT，里面有几个重要字段：</p>\n<pre><code class=\"language-json\">{\n  &quot;iss&quot;: &quot;https://accounts.example.com&quot;,\n  &quot;sub&quot;: &quot;acct_123&quot;,\n  &quot;aud&quot;: &quot;my-web-app&quot;,\n  &quot;exp&quot;: 1780000000,\n  &quot;email&quot;: &quot;user@example.com&quot;,\n  &quot;email_verified&quot;: true\n}\n</code></pre>\n<p>这里最重要的是三个：</p>\n<ul>\n<li><code>iss</code>：谁签发的身份。</li>\n<li><code>sub</code>：这个人在签发方那里的稳定身份 ID。</li>\n<li><code>aud</code>：这个身份结果是签给哪个应用的。</li>\n</ul>\n<p>应用拿到 <code>id_token</code> 后，不是看一眼邮箱就算登录成功。它要校验签名、校验 <code>iss</code>、校验 <code>aud</code>、校验过期时间，再用 <code>iss + sub</code> 找到或创建自己的本地用户。</p>\n<p>这件事看似麻烦，但它解决了一个非常根本的问题：身份不是靠前端说的，也不是靠 URL 里带的，更不是靠某个应用自己猜的。身份由一个明确的签发方签发，应用按标准验证。</p>\n<p>这就是边界。</p>\n<h2>为什么不是自己写一套登录接口</h2>\n<p>小应用一开始最容易这么做：</p>\n<pre><code class=\"language-text\">POST /login\n  email + code\n  -&gt; 返回 token\n</code></pre>\n<p>这当然能用。我也不反对在早期这么做。很多项目不是死在“不够标准”，而是死在还没上线就先给自己造了一座大教堂。</p>\n<p>问题在于，应用一多，这套简单接口会慢慢变形。</p>\n<p>第一个应用要邮箱登录。第二个应用要 GitHub 登录。第三个应用想接 Google。第四个应用又需要小程序登录。后台也要登录。CLI 也要登录。某个独立应用还想复用同一套账号。</p>\n<p>如果每个应用都自己做一遍，很快就会出现这些问题：</p>\n<ul>\n<li>用户在 A 应用登录了，到 B 应用又要重新注册。</li>\n<li>同一个邮箱在不同应用里变成不同用户。</li>\n<li>三方登录的回调、密钥、状态校验散在各处。</li>\n<li>邮箱验证码、限流、风控、封禁、审计重复建设。</li>\n<li>后面想做统一会员、跨应用权益、账号合并时，发现地基是散的。</li>\n</ul>\n<p>登录接口本身不复杂，复杂的是长期维护身份关系。</p>\n<p>OIDC 的意义就在这里。它不是让登录按钮变得更漂亮，而是把“谁来证明用户身份”这件事独立出来。</p>\n<p>一个中心身份服务负责认证：用户是谁、用什么方式登录、邮箱有没有验证、账号是否被禁用。各个应用负责业务：这个用户在本应用里叫什么、有什么角色、有没有套餐、额度还剩多少、是否欠费。</p>\n<p>这两个东西不要揉在一起。</p>\n<h2>中心身份和本地用户表</h2>\n<p>做统一登录时，很容易走向另一个极端：既然有统一身份中心，那是不是所有应用都共用一张用户表？</p>\n<p>我的结论是，不要。</p>\n<p>更稳的结构是：</p>\n<pre><code class=\"language-text\">统一身份中心\n  identity: acct_123\n  email: user@example.com\n  providers: email / google / github\n\n应用 A\n  user_id: 1\n  auth_issuer: https://accounts.example.com\n  auth_subject: acct_123\n  role: admin\n  plan: pro\n\n应用 B\n  user_id: 58\n  auth_issuer: https://accounts.example.com\n  auth_subject: acct_123\n  credits: 1200\n  status: active\n</code></pre>\n<p>中心身份回答“这是同一个人”。应用本地用户表回答“这个人在我这里是什么状态”。</p>\n<p>这个分工很重要。</p>\n<p>登录是全局问题，付费和权限往往是应用问题。一个人在写作工具里是会员，不代表他在图片工具里也有同样额度；一个人在后台里是管理员，不代表他在另一个产品里也该有管理权限；一个账号被中心禁用，所有应用都应该拦住，但某个应用里的业务封禁，也不一定要影响他使用别的应用。</p>\n<p>如果把所有东西都塞进 <code>id_token</code>，最后 token 会变成一个小型数据库。它看起来很强，实际很危险。</p>\n<p><code>id_token</code> 适合放身份事实，不适合放频繁变化的业务状态。套餐、额度、积分、订单、角色、封禁原因，这些最好还是在应用自己的系统里查。</p>\n<p><img src=\"/assets/posts/2026/oidc-login-identity/03-local-users-and-billing.jpg\" alt=\"一个中心身份映射到多个应用自己的用户账本\"></p>\n<h2>一个最小登录流程</h2>\n<p>OIDC 听起来有很多术语，但最常用的流程可以先按这一条理解：Authorization Code + PKCE。</p>\n<p>大概是这样：</p>\n<pre><code class=\"language-text\">1. 用户在应用里点击登录。\n2. 应用生成 state、nonce、code_verifier 和 code_challenge。\n3. 应用把用户跳转到身份中心 /authorize。\n4. 用户在身份中心完成登录。\n5. 身份中心带着临时 code 跳回应用。\n6. 应用校验 state。\n7. 应用用 code + code_verifier 去 /token 换 token。\n8. 应用校验 id_token。\n9. 应用用 iss + sub 查找或创建本地用户。\n10. 应用建立自己的 session。\n</code></pre>\n<p>这里有几个词容易混：</p>\n<p><code>code</code> 是一次性的临时票据，不是登录态。</p>\n<p><code>id_token</code> 是身份结果，用来证明用户是谁。</p>\n<p><code>access_token</code> 是访问资源的凭证，不应该被简单当成“用户是谁”的证明。</p>\n<p><code>refresh_token</code> 是续期凭证，能不能发、发给谁、怎么轮换，要非常谨慎。</p>\n<p>对普通 Web 应用来说，最舒服的结构通常是：后端完成回调和 token 交换，后端校验 <code>id_token</code>，然后给浏览器写自己的 HttpOnly Cookie。浏览器不需要长期拿着身份供应商的 token 在前端跑来跑去。</p>\n<p><img src=\"/assets/posts/2026/oidc-login-identity/02-code-token-flow.jpg\" alt=\"OIDC 登录中 code 回跳和后端 token 交换的简化路径\"></p>\n<p>一个极简的伪代码是：</p>\n<pre><code class=\"language-js\">app.get('/auth/start', (req, res) =&gt; {\n  const state = randomString();\n  const nonce = randomString();\n  const verifier = randomString();\n  const challenge = sha256base64url(verifier);\n\n  saveTemporaryLoginState(req, { state, nonce, verifier });\n\n  res.redirect(\n    'https://accounts.example.com/authorize?' +\n      new URLSearchParams({\n        response_type: 'code',\n        client_id: 'my-web-app',\n        redirect_uri: 'https://app.example.com/auth/callback',\n        scope: 'openid email profile',\n        state,\n        nonce,\n        code_challenge: challenge,\n        code_challenge_method: 'S256',\n      }),\n  );\n});\n\napp.get('/auth/callback', async (req, res) =&gt; {\n  const saved = readTemporaryLoginState(req);\n  if (req.query.state !== saved.state) {\n    return res.status(400).send('bad state');\n  }\n\n  const tokens = await exchangeCodeForTokens({\n    code: req.query.code,\n    codeVerifier: saved.verifier,\n  });\n\n  const claims = await verifyIdToken(tokens.id_token, {\n    issuer: 'https://accounts.example.com',\n    audience: 'my-web-app',\n    nonce: saved.nonce,\n  });\n\n  const user = await findOrCreateLocalUser({\n    issuer: claims.iss,\n    subject: claims.sub,\n    email: claims.email,\n  });\n\n  writeLocalSessionCookie(res, user.id);\n  res.redirect('/dashboard');\n});\n</code></pre>\n<p>真实项目里还要处理错误、过期、回跳地址白名单、Cookie 安全属性、日志和限流。但主干就是这几步。</p>\n<h2>什么场景适合用 OIDC</h2>\n<p>OIDC 适合这些场景：</p>\n<p>第一，有多个应用需要共用登录。</p>\n<p>哪怕现在只有两个应用，也值得提前想一下。登录一旦散出去，后面再收回来会很费劲。用户数据、三方账号、邮箱验证、旧 token、历史订单，全会变成迁移问题。</p>\n<p>第二，想把“登录”和“业务用户”分开。</p>\n<p>这对有付费、积分、额度、团队、角色、封禁等业务状态的应用尤其重要。中心身份不要变成超级业务用户表，应用本地用户表也不要重复造一整套登录供应商。</p>\n<p>第三，需要接入第三方身份。</p>\n<p>Google、Microsoft、企业身份、组织账号，这些天然适合按 OIDC 思路理解。即使某些平台表面文档写的是 OAuth 登录，背后也往往会给 <code>id_token</code> 或提供类似的身份验证能力。</p>\n<p>第四，需要给 CLI、移动端、独立服务提供统一登录。</p>\n<p>Web 可以走 Authorization Code + PKCE。CLI 在有浏览器时也可以走 PKCE；没有浏览器时，可以考虑 device code flow。移动端要用系统浏览器或平台推荐方式，不要让用户在不可信 WebView 里输入密码。</p>\n<p>第五，未来可能做跨应用会员或统一账号治理。</p>\n<p>付费本身不一定在身份中心里做，但身份中心至少要稳定回答“这是哪个人”。否则订单、权益、额度、风控很难跨应用协调。</p>\n<h2>什么场景不必急着用</h2>\n<p>不是所有项目都需要 OIDC。</p>\n<p>如果只是一个很小的团队工具，只有一个 Web 端，用户也很少，简单 session 登录足够。</p>\n<p>如果应用完全依赖某个平台，比如只在某个小程序里运行，用户体系也只围绕平台 openid 展开，那么先把平台登录做好更实际。</p>\n<p>如果只是服务间调用，机器访问机器，重点不是“用户是谁”，而是“这个服务有没有权限”，那 API key、mTLS、OAuth client credentials 可能更合适。</p>\n<p>如果只是想让用户不用输密码，WebAuthn / Passkey 是很好的登录方式，但它不是 OIDC 的替代品。Passkey 更像一种认证手段，OIDC 更像多个应用之间传递身份结果的协议。</p>\n<p>还有一种情况也要谨慎：为了显得正规，先上一个很重的身份平台，然后业务还没跑起来，配置、回调、证书、client、realm、scope 已经把人绕晕了。</p>\n<p>标准是为了降低长期复杂度，不是为了在第一天制造复杂度。</p>\n<h2>和几种常见方案的区别</h2>\n<p>OAuth 2.0 和 OIDC 最容易混。</p>\n<p>OAuth 的核心是授权。比如一个应用想读取用户的网盘文件，用户同意以后，应用拿到 access token 去访问网盘 API。这里重点是“能不能访问资源”。</p>\n<p>OIDC 在 OAuth 上加了身份层。它标准化了 <code>id_token</code>、<code>sub</code>、<code>iss</code>、<code>aud</code>、<code>nonce</code>、Discovery、JWKS、UserInfo。它让应用可以按标准确认“登录用户是谁”。</p>\n<p>SAML 更老，也很常见，尤其在企业 SSO 里。它基于 XML，企业软件生态成熟，但对现代 Web 和移动应用来说，开发体验通常不如 OIDC 轻。</p>\n<p>CAS 多见于高校和早期组织系统，单点登录能力很直接，但生态和现代 API 场景不如 OIDC 普遍。</p>\n<p>LDAP / Active Directory 更像目录服务。它擅长保存组织、用户、组和凭据，也常用于企业目录认证。但直接让现代 Web 应用到处接 LDAP，通常不如先放一层 OIDC Provider。</p>\n<p>Session Cookie 是单个 Web 应用最常见的登录态。它和 OIDC 不冲突。很多时候最稳的做法正是：外部登录用 OIDC，应用自己的登录态用 HttpOnly Session Cookie。</p>\n<p>JWT 也不是 OIDC 的替代品。JWT 只是 token 格式。自己签一个 JWT 不等于实现了 OIDC。OIDC 关心的是签发方、受众、发现机制、公钥、流程和校验规则。</p>\n<p>WebAuthn / Passkey 解决的是“用户如何证明自己”。OIDC 解决的是“应用如何相信一个身份供应商给出的身份结果”。它们可以一起用：用户用 Passkey 登录身份中心，应用通过 OIDC 接收身份结果。</p>\n<h2>AI 时代为什么更该懂这些概念</h2>\n<p>AI 让写代码变快了，但也让很多系统问题更早暴露。</p>\n<p>以前一个应用可以靠一张用户表和几个接口撑很久。现在应用之间更容易互相连接，后台、CLI、Agent、移动端、独立服务、自动化任务都可能变成同一套产品能力的入口。</p>\n<p>AI 也会让更多人参与构建系统。一个人可以同时做前端、后端、后台、部署、支付、登录。能力变强以后，最危险的不是写不出代码，而是不知道边界应该放在哪里。</p>\n<p>登录尤其如此。</p>\n<p>登录一旦做错，后面不是改几个页面那么简单。用户身份、付费记录、权限、风控、账号合并、注销、审计，全都绑在一起。代码可以让 AI 帮忙写，概念要自己想明白。</p>\n<p>我现在越来越觉得，独立开发者和小团队至少要有几个底层判断：</p>\n<ul>\n<li>登录不是一个按钮，是身份边界。</li>\n<li>付费不是一个回调，是业务权益和账本。</li>\n<li>中心身份不要吞掉应用用户表。</li>\n<li>应用用户表也不要重复造身份中心。</li>\n<li>token 不是越多越好，scope 不是越大越好。</li>\n<li>用户是谁，和用户能做什么，要分开建模。</li>\n</ul>\n<p>这些判断不花哨，但很管用。</p>\n<h2>最后</h2>\n<p>OIDC 不是必须每个项目第一天就上的东西。</p>\n<p>但只要应用开始变多，只要你希望用户在多个产品之间有同一个身份，只要你不想在每个应用里重复维护邮箱、三方登录、验证码、风控和账号合并，就应该认真理解它。</p>\n<p>它解决的不是“怎么做一个登录页”，而是“谁有资格证明用户是谁”。</p>\n<p>这个问题一旦想清楚，后面的很多设计都会顺一点：身份中心负责认证，应用负责业务；<code>id_token</code> 证明身份，本地 session 维持登录；<code>iss + sub</code> 做稳定映射，付费和权限留在应用自己的账本里。</p>\n<p>技术协议最后都会落回这种普通道理。</p>\n<p>门要有人看，账要有人记。一个应用若想长期活下去，登录和付费这两件事，总归不能糊涂。</p>\n","date_published":"2026-06-20T00:00:00.000Z","tags":["OIDC","OAuth","登录","身份认证","架构"],"language":"zh"},{"id":"https://www.lihuanyu.com/posts/ai-toolchain-performance-matters-in-ai-era/","url":"https://www.lihuanyu.com/posts/ai-toolchain-performance-matters-in-ai-era/","title":"AI 时代，工具链性能变得前所未有的重要","summary":"AI 编程真正改变的不是“写代码更快”这么简单，而是开发进入了 agentic loop：AI 写代码、跑验证、读错误、再修复。这个循环里，test、lint、build 的耗时会直接决定整体吞吐。","content_html":"<p>D2 现场听尤雨溪分享时，有一张图让我印象很深。</p>\n<p>它不是那种靠视觉效果取胜的发布会大图，更像是把一个大家已经隐约感觉到的变化，直接摆到了台面上：AI 把写代码这件事加速以后，开发循环里最慢的地方变了。</p>\n<p>图里有三个阶段。</p>\n<p>AI 之前，大部分时间都花在 writing code 上，waiting for tools 只是下面一小块。人写代码本来就慢，工具链慢一些，当然烦，但不一定是最刺眼的问题。</p>\n<p>AI 之后，writing code 被压缩了。代码出来得很快，waiting for tools 那一块反而变得很显眼。</p>\n<p>再往后，是 AI + better tools：更快的基本错误反馈，更快的行为验证，更快的迭代速度。</p>\n<p><img src=\"/assets/posts/2026/ai-toolchain-performance/01-d2-ai-era-toolchain-comparison.jpg\" alt=\"D2 现场图展示 AI 时代前后写代码和等待工具链反馈的比例变化\"></p>\n<p>这张图讲中的不是“开发者等命令很烦”这种老问题，而是一个新的工程事实：AI 时代，工具链性能会直接决定 AI 编程的循环速度。</p>\n<p><a href=\"/en/posts/2026/toolchain-performance-matters-more-than-ever-in-the-ai-era/\">English version: Toolchain Performance Matters More Than Ever in the AI Era</a></p>\n<h2>不是人在等，是循环在等</h2>\n<p>过去说工具链慢，通常是在说人等得烦。</p>\n<p>跑测试，去喝水；跑构建，切出去看一眼消息；lint 卡一下，顺手打开别的窗口。慢当然不舒服，但它很多时候只是人的工作流里一段空白。</p>\n<p>现在 Codex、Claude Code、Cursor 这类工具把开发方式改成了另一种形态。</p>\n<p>它们不是只生成一段代码就结束。更典型的过程是：读上下文，改代码，跑测试或 lint，读取错误，再改，再跑。这个过程越来越像一个 agentic loop。</p>\n<p>在这个循环里，工具链不是“人类开发者旁边的辅助命令”，而是 AI 判断下一步的输入来源。测试、lint、构建给出的反馈，正在变成它的感官系统。</p>\n<p>测试结果返回慢，AI 就慢。</p>\n<p>lint 返回慢，AI 就慢。</p>\n<p>构建返回慢，AI 也慢。</p>\n<p>所以问题不再是“人有没有耐心等 10 分钟”。问题变成了：一次自动开发循环的反馈延迟是多少？一个 agent 在同样时间里能完成多少轮验证和修正？</p>\n<p>以前工具链慢，是人在等。</p>\n<p>现在工具链慢，是整个开发循环在等。</p>\n<p><img src=\"/assets/posts/2026/ai-toolchain-performance/02-ai-feedback-bottleneck.jpg\" alt=\"AI 辅助写代码后，验证环节变成更明显的反馈瓶颈\"></p>\n<h2>为什么以前没这么痛</h2>\n<p>以前人写代码，慢是默认值。</p>\n<p>一个功能从理解需求到写完，可能要半小时、两小时，甚至一天。中间跑几次测试，每次慢一点，当然烦，但总耗时里还有大量人工思考、查资料、改代码、对齐上下文。</p>\n<p>工具链的慢，被人的慢盖住了。</p>\n<p>这不是说工具链性能以前不重要。只是它没有那么容易成为主要矛盾。</p>\n<p>AI 把这个比例改了。</p>\n<p>当一个改动可以很快被生成出来，真正决定节奏的就不是“代码出现得快不快”，而是“代码被验证得快不快”。如果 AI 30 秒写完，测试要跑 10 分钟，那么整个循环的速度并不由 AI 决定，而由测试决定。</p>\n<p>这有点像把工厂里最慢的一段传送带暴露出来。</p>\n<p>以前前面的工人动作慢，后面的传送带慢一点没人注意。现在前面换成了高速机械臂，传送带的速度就成了产线速度。</p>\n<h2>统一工具链为什么变得更有意义</h2>\n<p>尤雨溪在 D2 上讲 Oxc、Rolldown、Vite、Vitest，我理解它不只是“新工具更快”。</p>\n<p>更重要的是，JS 工具链过去太碎了。这种碎，在人类开发节奏下可以忍；在 agentic loop 里，会被放大成系统成本。</p>\n<p>parser、transformer、test runner、linter、formatter、bundler，各自有各自的实现，各自 parse 一遍，各自 transform 一遍，各自配置一遍。很多时间不是花在业务验证上，而是花在重复处理同一批代码上。</p>\n<p>AI 时代，这种重复开销会被放大。</p>\n<p>因为 AI 不会像人一样一天只认真跑几次验证。一个 agent 在修一个问题时，可能会非常自然地多次跑测试、多次跑 lint、多次做局部验证。它不觉得“麻烦”，它只是按循环推进。</p>\n<p>如果工具链足够快，这种频繁验证会带来更高质量的代码。</p>\n<p>如果工具链很慢，频繁验证就会把循环拖死。</p>\n<p>所以 Oxc、Rolldown、Vite、Vitest 这类工具链的意义，不只是 benchmark 数字漂亮。它们是在把 AI 编程所需的反馈回路变短。</p>\n<p>这也是为什么前端工具链的变化，对后端项目同样有启发。</p>\n<p>语言、框架、业务形态可以不同，但 agentic loop 的结构是一样的：改代码，验证行为，读反馈，继续修正。慢的验证链路，会限制整个系统。</p>\n<h2>一个后端项目里的数据点</h2>\n<p>观点说到这里，需要一个足够具体的数字。</p>\n<p>手边有个中等规模的 NestJS 后端项目，测试文件已经积累到 200 个左右。</p>\n<p>迁移前，单测跑在 Jest 上。Jest 成熟、稳定、资料多，很多 NestJS 项目默认就是它。但随着测试数量增加，反馈时间开始变得很难忍。</p>\n<p>一次完整单测跑到 580 多秒还没有结束。</p>\n<p>这里要说清楚：不是 580 秒跑完，而是超过 580 秒仍未完成。所以这个数字只能作为旧流程耗时的下界。</p>\n<p>迁到 Vitest + SWC 后，当时同一批单测完整通过：</p>\n<pre><code class=\"language-text\">Test Files  200 passed (200)\nTests       1299 passed (1299)\nDuration    34.73s\n</code></pre>\n<p>只按这个下界算，也至少是 16.7 倍差距。真实差距大概率更高，因为旧流程没有等到结束。</p>\n<p>后来项目又多了一些测试，当前版本复跑结果是：</p>\n<pre><code class=\"language-text\">Test Files  206 passed (206)\nTests       1314 passed (1314)\nDuration    37.27s\nshell total 38.345s\n</code></pre>\n<p>这个变化对 agentic loop 很关键。它不是一次工具迁移的“成绩单”，而是一个反馈回路变短后的样子。</p>\n<p>30 多秒跑完整单测，AI 可以把它当成常规验证步骤。超过 10 分钟还没结束，AI 每一轮修正都会被迫停在原地。</p>\n<p><img src=\"/assets/posts/2026/ai-toolchain-performance/03-test-time-comparison.jpg\" alt=\"单测耗时从超过 580 秒仍未完成，变成 34.73 秒完整通过\"></p>\n<h2>为什么我会先动 test</h2>\n<p>如果只能先优化一个环节，我会先看 test。</p>\n<p>原因不是 lint 不重要，而是 test 在 agentic loop 里提供的信息密度更高。</p>\n<p>lint 主要告诉你规则、格式、一部分静态问题有没有踩线。test 更接近行为验证：service 的分支有没有断，controller 的返回有没有变，guard 的边界有没有破，provider 之间的契约还能不能成立。</p>\n<p>AI 写代码时，最需要的不是“看起来格式正确”，而是“这个改动有没有破坏已有行为”。</p>\n<p>所以单测如果很慢，AI 的修复循环会明显变钝。它可能能改，但每一次判断都要等很久。等得久，循环次数就少；循环次数少，能自我修正的空间就小。</p>\n<p>Vitest + SWC 的收益就在这里。</p>\n<p>它不是把命令名字从 <code>jest</code> 换成 <code>vitest</code>。它把“完整验证”从一个需要斟酌的动作，变成了循环里可以频繁发生的动作。</p>\n<p>Oxlint 也是同一个方向。</p>\n<p>lint 切到 Oxlint 后，日常 <code>pnpm lint</code> 基本进入秒级，甚至有时亚秒级完成。这里我不硬写倍数，因为没有保留迁移前 ESLint 的严谨基线。可以确定的是，它从“需要等一下”变成了“可以顺手跑一下”。</p>\n<p>在 AI 编程里，这个差别很大。</p>\n<p><img src=\"/assets/posts/2026/ai-toolchain-performance/04-nestjs-validation-loop.jpg\" alt=\"NestJS 后端项目里的 test、lint、build 和 CI 反馈环路\"></p>\n<h2>快不是为了爽，是为了多跑几轮</h2>\n<p>工具链性能经常被说成“开发体验”。</p>\n<p>这当然没错。命令快一点，开发者舒服一点，没人反对。</p>\n<p>但在 AI 时代，只把它理解成体验还不够。</p>\n<p>更快的工具链会改变开发策略。</p>\n<p>测试跑 30 秒，你会倾向于让 AI 改完就跑完整测试。测试跑 10 分钟，你会倾向于只跑相关文件，甚至先不跑。lint 亚秒级结束，它就会变成每次修改后的自然步骤。lint 要等很久，它就会被推迟。</p>\n<p>工具链慢，会让正确动作变贵。</p>\n<p>工具链快，会让正确动作重新便宜。</p>\n<p>对 agentic loop 来说，“便宜”很重要。因为 agent 不怕重复，它怕反馈慢。反馈越快，它越能多尝试、多修正、多收敛。反馈越慢，它越容易卡在单轮修改里。</p>\n<p>这也是我现在看工具链性能的方式。</p>\n<p>它不是锦上添花，而是在决定 AI 编程的吞吐上限。</p>\n<h2>不要把迁移写成追新工具</h2>\n<p>迁 Vitest、Oxlint，或者将来看 Oxc、Rolldown，不应该只是因为它们新。</p>\n<p>追新工具很容易变成另一种折腾。</p>\n<p>真正值得迁的理由，是它们能不能缩短反馈回路，能不能减少重复工作，能不能让 agentic loop 更稳定地跑起来。</p>\n<p>这也是“一刀切”有时反而更合理的原因。</p>\n<p>如果 Jest 和 Vitest 长期并存，短期看起来温和，长期会变成两套心智负担。mock API、globals、配置、覆盖率、watch 行为、类型声明，都可能出现微妙差异。AI 修改测试时，也更容易混用两套写法。</p>\n<p>工具链基础设施最怕半迁移。</p>\n<p>半迁移不是保守，很多时候只是把判断成本留给未来每一次修改。人要判断，AI 也要判断。判断多了，循环就慢，错误也会变多。</p>\n<p>当然，一刀切不是闭眼替换。常用 mock、fake timers、覆盖率、e2e、integration 的边界都要确认。但确认之后，统一工具链本身就是收益。</p>\n<h2>最后的判断</h2>\n<p>AI 时代，工具链性能变得前所未有的重要。</p>\n<p>不是因为大家突然都成了性能洁癖，也不是因为新工具名字更酷。真正的原因很朴素：AI 把代码生成速度大幅提高以后，验证链路就成了新的主要矛盾。</p>\n<p>更准确地说，是 agentic loop 的主要矛盾。</p>\n<p>AI 写代码，跑验证，读错误，再修代码。这个循环能不能快，取决于每一轮反馈能不能快。test、lint、build 不再只是命令行里的几条命令。</p>\n<p>从 Jest 迁到 Vitest + SWC 的数据只是一个佐证：一个超过 580 秒仍未完成的单测流程，变成了 34.73 秒完整通过 200 个测试文件、1299 个测试用例。这个变化真正改变的，是循环速度。</p>\n<p>工具链不应该成为 AI 编程里的红灯。</p>\n<p>它应该是一条短而清楚的反馈回路：改一点，验一下，错了马上修，对了继续走。</p>\n<p>写代码只是开始。</p>\n<p>能快速验证代码还对，才是 AI 时代工程效率的底座。</p>\n","date_published":"2026-06-14T00:00:00.000Z","tags":["AI","工程化","Vitest","Oxc","NestJS"],"language":"zh"},{"id":"https://www.lihuanyu.com/en/posts/2026/toolchain-performance-matters-more-than-ever-in-the-ai-era/","url":"https://www.lihuanyu.com/en/posts/2026/toolchain-performance-matters-more-than-ever-in-the-ai-era/","title":"Toolchain Performance Matters More Than Ever in the AI Era","summary":"AI coding is increasingly an agentic loop: write code, run validation, read failures, and fix again. In that loop, test, lint, and build performance directly shape engineering throughput.","content_html":"<p>At D2, while listening to Evan You’s talk in person, one slide stayed with me.</p>\n<p>It was not a flashy conference slide. It was more like a clean way to name something many of us had already started to feel: once AI accelerates the act of writing code, the slowest part of the development loop moves somewhere else.</p>\n<p>It shows three stages.</p>\n<p>Before AI, most of the time was spent writing code. Waiting for tools was only a small block near the bottom. Humans were already slow enough that slow tooling was annoying, but not always the most visible constraint.</p>\n<p>After AI, the writing-code block gets compressed. Code appears much faster, and waiting for tools suddenly takes a much larger share of the loop.</p>\n<p>Then comes AI + better tools: faster basic error feedback, faster behavior validation, and faster iteration.</p>\n<p><img src=\"/assets/posts/2026/ai-toolchain-performance/01-d2-ai-era-toolchain-comparison.jpg\" alt=\"A D2 talk photo showing how the proportion between writing code and waiting for tooling changes before and after AI\"></p>\n<p>That photo is not only about the old annoyance of developers waiting for commands. It points at a newer engineering fact: in the AI era, toolchain performance directly affects the speed of AI-assisted development loops.</p>\n<p><a href=\"/posts/ai-toolchain-performance-matters-in-ai-era/\">Chinese version of this article</a></p>\n<h2>The Loop Is Waiting</h2>\n<p>When we used to say tooling was slow, we usually meant that humans had to wait.</p>\n<p>Run tests, get water. Run a build, check messages. Wait for lint, switch to another window. It was unpleasant, but it was still just one empty patch inside a human workflow.</p>\n<p>Tools like Codex, Claude Code, and Cursor have changed the shape of development.</p>\n<p>They do not only generate a piece of code and stop. A more typical process is: read context, edit code, run tests or lint, read the failure, edit again, run again. The workflow is becoming an agentic loop.</p>\n<p>In that loop, the toolchain is not a side command sitting next to a human developer. It is the input stream that lets the agent decide what to do next. Tests, lint, and builds are becoming its sensory system.</p>\n<p>Slow test feedback makes the agent slower.</p>\n<p>Slow lint feedback makes the agent slower.</p>\n<p>Slow build feedback makes the agent slower.</p>\n<p>So the question is no longer whether a human has the patience to wait ten minutes. The question is: what is the feedback latency of one automated development loop? How many validation and correction cycles can an agent complete in the same amount of time?</p>\n<p>Before, slow tooling meant humans were waiting.</p>\n<p>Now, slow tooling means the whole development loop is waiting.</p>\n<p><img src=\"/assets/posts/2026/ai-toolchain-performance/02-ai-feedback-bottleneck.jpg\" alt=\"After AI-assisted coding, validation becomes the visible feedback bottleneck\"></p>\n<h2>Why It Was Less Painful Before</h2>\n<p>Human coding used to be slow by default.</p>\n<p>A feature could take half an hour, two hours, or a full day. During that time, the developer was reading context, checking interfaces, thinking through edge cases, making careful edits, and occasionally running tests. Slow tooling was still irritating, but it was surrounded by a lot of human work.</p>\n<p>The slowness of the toolchain was hidden by the slowness of humans.</p>\n<p>That does not mean toolchain performance was unimportant before. It means it was less likely to become the primary constraint.</p>\n<p>AI changes that ratio.</p>\n<p>When a change can be generated quickly, the key question is no longer only how fast code appears. It becomes how fast the code can be validated. If AI writes a patch in 30 seconds and the test suite takes ten minutes, the loop is not governed by AI. It is governed by tests.</p>\n<p>It is like exposing the slowest belt in a factory line.</p>\n<p>When the worker at the front was slow, a slow conveyor belt later in the line was less noticeable. Replace the worker with a fast robotic arm, and the belt speed becomes the factory speed.</p>\n<h2>Why Unified Toolchains Matter More Now</h2>\n<p>When Evan You talks about Oxc, Rolldown, Vite, and Vitest, I do not read it as merely “new tools are faster.”</p>\n<p>The deeper point is that JavaScript tooling has been highly fragmented. That fragmentation was tolerable at human coding speed. Inside an agentic loop, it turns into system cost.</p>\n<p>Parsers, transformers, test runners, linters, formatters, and bundlers often have their own implementations. They parse the same code separately, transform it separately, configure behavior separately, and sometimes disagree in subtle ways. A lot of time is spent re-processing the same program rather than validating product behavior.</p>\n<p>In the AI era, that waste gets amplified.</p>\n<p>An agent does not treat repeated validation as a nuisance in the same way humans do. While fixing a problem, it may naturally run tests multiple times, run lint multiple times, and perform focused checks repeatedly. That is exactly what we want if the feedback is fast.</p>\n<p>If the toolchain is fast, frequent validation improves code quality.</p>\n<p>If the toolchain is slow, frequent validation stalls the loop.</p>\n<p>That is why tools such as Oxc, Rolldown, Vite, and Vitest matter beyond benchmark screenshots. They shorten the feedback loop that AI-assisted programming depends on.</p>\n<p>This is not only a frontend issue.</p>\n<p>My concrete migration happened in a NestJS backend project. The framework and runtime are different, but the loop has the same shape: change code, validate behavior, read feedback, and continue. A slow validation chain limits the whole system.</p>\n<h2>One Backend Data Point</h2>\n<p>This argument needs at least one concrete number.</p>\n<p>One medium-sized NestJS backend project had accumulated around 200 test files.</p>\n<p>Before the migration, unit tests ran on Jest. Jest is mature, stable, and well documented. Many NestJS projects started there for good reasons. But as the suite grew, the feedback time became painful.</p>\n<p>One full unit-test run went beyond 580 seconds without finishing.</p>\n<p>That distinction matters: it did not finish in 580 seconds. It was still running after more than 580 seconds, so that number is only a lower bound for the old flow.</p>\n<p>After moving to Vitest + SWC, the same batch completed:</p>\n<pre><code class=\"language-text\">Test Files  200 passed (200)\nTests       1299 passed (1299)\nDuration    34.73s\n</code></pre>\n<p>Even if the old run is treated only as a lower bound, that is at least a 16.7x difference. The real gap was probably larger because the old run never reached completion.</p>\n<p>The project later gained a few more tests. A current rerun looked like this:</p>\n<pre><code class=\"language-text\">Test Files  206 passed (206)\nTests       1314 passed (1314)\nDuration    37.27s\nshell total 38.345s\n</code></pre>\n<p>This matters a lot for an agentic loop. It is not the scorecard of a tool migration. It is what a shortened feedback loop looks like.</p>\n<p>When a full unit-test suite takes around half a minute, an agent can treat it as a normal validation step. When it has already exceeded ten minutes and is still running, every correction cycle is forced to stop in place.</p>\n<p><img src=\"/assets/posts/2026/ai-toolchain-performance/03-test-time-comparison.jpg\" alt=\"The unit test suite moved from more than 580 seconds and still unfinished to a completed 34.73-second run\"></p>\n<h2>Why I Would Start With Tests</h2>\n<p>If only one part can be optimized first, I would look at tests before lint.</p>\n<p>Not because lint is unimportant, but because tests carry denser feedback inside an agentic loop.</p>\n<p>Lint tells you whether formatting, rules, and some static constraints are fine. Tests get closer to behavior: whether a service branch broke, whether a controller response changed, whether a guard boundary still holds, whether provider contracts still work.</p>\n<p>AI coding needs more than “this looks syntactically acceptable.” It needs to know whether the change broke existing behavior.</p>\n<p>When tests are slow, the AI repair loop becomes blunt. It may still be able to fix the problem, but each judgment takes too long. Fewer cycles means less room for self-correction.</p>\n<p>That is where Vitest + SWC helped.</p>\n<p>The point was not changing a command name from <code>jest</code> to <code>vitest</code>. The point was turning full validation from a decision that needed hesitation into an action that could happen frequently inside the loop.</p>\n<p>Oxlint follows the same direction.</p>\n<p>After moving lint to Oxlint, daily <code>pnpm lint</code> runs are generally in the seconds range, and sometimes below a second. I do not want to attach a precise multiplier because I did not keep a rigorous ESLint baseline. The reliable statement is simpler: lint moved from “something that requires a wait” to “something you can run casually.”</p>\n<p>In AI-assisted development, that difference is large.</p>\n<p><img src=\"/assets/posts/2026/ai-toolchain-performance/04-nestjs-validation-loop.jpg\" alt=\"A privacy-safe NestJS backend feedback loop: test, lint, build, and CI validation\"></p>\n<h2>Fast Means More Cycles</h2>\n<p>Toolchain performance is often described as developer experience.</p>\n<p>That is true. Faster commands feel better. No one argues with that.</p>\n<p>But in the AI era, experience is not the whole story.</p>\n<p>Faster tooling changes the development strategy.</p>\n<p>If tests take 30 seconds, you are more likely to let the agent run the full suite after a meaningful change. If tests take ten minutes, you are more likely to run only related files, or postpone validation. If lint finishes in less than a second, it becomes a natural step after edits. If lint takes a long time, it gets delayed.</p>\n<p>Slow tooling makes correct behavior expensive.</p>\n<p>Fast tooling makes correct behavior cheap again.</p>\n<p>Cheap matters to agentic loops. Agents are not afraid of repetition. They are constrained by feedback latency. The faster the feedback, the more they can try, fix, and converge. The slower the feedback, the more they get stuck in single-shot edits.</p>\n<p>That is how I now think about toolchain performance.</p>\n<p>It is not polish. It is a throughput limit for AI programming.</p>\n<h2>This Is Not Tool-Chasing</h2>\n<p>Moving to Vitest, Oxlint, or eventually looking at Oxc and Rolldown should not be about chasing whatever is new.</p>\n<p>That kind of tool-chasing easily becomes churn.</p>\n<p>The real reason to migrate is whether the tool shortens the feedback loop, reduces repeated work, and makes the agentic loop more stable.</p>\n<p>This is also why a clean cut can sometimes be better.</p>\n<p>Keeping Jest and Vitest side by side may look safer in the short term, but it can create long-term cognitive overhead. Mock APIs, globals, configuration, coverage behavior, watch mode, and type declarations may all differ slightly. AI-generated test changes are also more likely to mix styles when both systems remain present.</p>\n<p>Half-migrated infrastructure is expensive.</p>\n<p>It leaves judgment cost for every future change. Humans have to decide. AI has to decide. More decisions mean more latency and more mistakes.</p>\n<p>A clean cut does not mean replacing things blindly. Mocks, fake timers, coverage, e2e, and integration boundaries still need to be checked. But once those checks are done, a unified toolchain is itself a benefit.</p>\n<h2>The Practical Conclusion</h2>\n<p>Toolchain performance matters more than ever in the AI era.</p>\n<p>Not because everyone suddenly became obsessed with performance, and not because new tool names sound exciting. The reason is simpler: once AI dramatically accelerates code generation, validation becomes the new primary constraint.</p>\n<p>More precisely, it becomes the primary constraint of the agentic loop.</p>\n<p>AI writes code, runs validation, reads failures, and edits again. The speed of that loop depends on the speed of feedback. Tests, lint, and build are no longer just command-line chores.</p>\n<p>The Jest to Vitest + SWC migration is only one data point: a unit-test flow that had gone beyond 580 seconds without finishing became a 34.73-second completed run across 200 test files and 1,299 test cases. What changed was the loop speed.</p>\n<p>The toolchain should not be the red light in AI programming.</p>\n<p>It should be a short, clear feedback loop: change a little, verify, fix immediately if needed, and continue.</p>\n<p>Writing code is only the beginning.</p>\n<p>Being able to quickly prove that the code still works is the foundation of engineering efficiency in the AI era.</p>\n","date_published":"2026-06-14T00:00:00.000Z","tags":["AI","Engineering","Vitest","Oxc","NestJS"],"language":"en"},{"id":"https://www.lihuanyu.com/en/posts/2026/the-next-step-for-admin-platforms-is-cli-and-skill/","url":"https://www.lihuanyu.com/en/posts/2026/the-next-step-for-admin-platforms-is-cli-and-skill/","title":"The Next Step for Admin Platforms Is CLI and Skill","summary":"A reflection on why admin and configuration platforms should stop exposing their capabilities only through pages, and start turning operations into CLI commands and agent-readable skills.","content_html":"<p>I have been getting a little tired of admin forms lately.</p>\n<p><img src=\"/assets/posts/2026/admin-cli-skill/01-admin-console-cli.jpg\" alt=\"An admin form and CLI interface living side by side\"></p>\n<p>Not because one particular page is terrible, and not because some component library invented another strange interaction. The annoying part is something else: many admin systems look like they have a lot of modules, but once you peel them apart, they repeat the same kind of action.</p>\n<p>Open a list. Click create. Fill in name, time, scope, priority, budget, material, target URL. Save. Submit for approval. Wait for it to take effect. Switch to another business module, and the fields change, but the flow does not change much. Switch again to a technical configuration platform, and it is still close: change a flag, tune a threshold, pick an environment, inspect the diff, publish, rollback.</p>\n<p>Of course admin platforms are not only forms. They have lists, charts, approval flows, permissions, imports, exports, logs, and dashboards. But the high-frequency work often returns to a few plain actions:</p>\n<p>Read some data. Change some configuration. Trigger a task.</p>\n<p>To say it more directly, many admin platforms are interfaces wrapped around APIs so that humans can operate them.</p>\n<p>That wrapper used to be necessary. You cannot ask operations people to change campaigns with curl. You cannot ask customer support to connect to a database and update status fields. You cannot ask business users to remember dozens of API parameters. Pages block risk, spread fields out, add permissions and validation, and at least let people know what they are doing.</p>\n<p>But after AI agents appeared, that assumption started to shift.</p>\n<p>Pages still matter. But pages should not be the only entrance to a platform’s capabilities.</p>\n<p><a href=\"/posts/2026/%E4%B8%AD%E5%90%8E%E5%8F%B0%E7%9A%84%E4%B8%8B%E4%B8%80%E7%AB%99%E6%98%AFCLI%E5%92%8CSkill/\">Chinese version of this article</a></p>\n<h2>Admin Pages Protect People</h2>\n<p>When I built admin systems in the past, I was fairly tolerant of complex forms.</p>\n<p>A campaign configuration page with dozens of fields is not necessarily a product manager committing a crime. Campaigns really do have time windows, inventory, channels, audiences, materials, budget, risk control, gradual rollout, and fallback behavior. If the page does not lay these things out, people are more likely to cause incidents.</p>\n<p>Many verbose admin pages are actually protecting the operator.</p>\n<p>They encode rules such as “this field cannot be empty”, “this city cannot be targeted”, “this budget requires approval”, and “this slot cannot carry two campaigns at the same time” into the page and backend validation. Operations may move more slowly, but slow is still better than a misconfigured production campaign.</p>\n<p>The problem is that this design assumes one thing: every detail has to be completed by a human, step by step.</p>\n<p>Before AI, that default made sense. If humans do not fill in the forms, who will? If humans do not click publish, who will? But agents can now read documents, inspect APIs, generate parameters, run commands, check results, and then ask people to confirm. Keeping every capability hidden behind pages begins to feel wasteful.</p>\n<p>What is worse, asking an agent to operate a GUI is often awkward.</p>\n<p>Where is the button? Did the modal open? Which page of the table is visible? Is the field collapsed? These are interface details for humans, but noise for agents. An agent does not need a polished sidebar or an 8px rounded button. It needs a stable entrance:</p>\n<pre><code class=\"language-bash\">ops campaign preview --id 123 --format json\nops campaign publish --id 123 --dry-run\nops campaign rollback --id 123 --reason &quot;inventory config error&quot;\n</code></pre>\n<p>These commands are not pretty, but they are readable by machines and inspectable by humans.</p>\n<h2>CLI Fits Agents Better Than Pages</h2>\n<p>Developer tools have already demonstrated this many times.</p>\n<p><code>git status</code>, <code>pnpm test</code>, <code>kubectl get pods -o json</code>, and <code>gh pr create</code> do not have much product packaging, but agents use them smoothly. The reason is simple: input is text, output is controllable, failure has an exit code, the process can be recorded, and commands can be chained.</p>\n<p>If an admin platform is only for humans, GUI is the natural choice.</p>\n<p>But if the same platform also needs to serve agents, CLI is hard to avoid.</p>\n<p>The benefit of CLI is not that it is sophisticated. The benefit is that it is plain enough. It can be called by a shell, by CI, by scripts, by a logging system, and reproduced locally. When something goes wrong, people do not have to remember “which button did I click?” They can look at the command, the parameters, the request id, and the approval record.</p>\n<p>That matters for admin systems.</p>\n<p>Admin operations are usually not toys. One configuration may affect campaign revenue. One flag may change a user path. One policy threshold may touch risk control. Agents can execute, but the execution must leave a trace. CLI is much more reliable than GUI automation for that job.</p>\n<p>So I increasingly think admin platforms should have two entrances.</p>\n<p>GUI for people to inspect. CLI for agents to run.</p>\n<p>Not replacement. Division of labor.</p>\n<h2>Operators Should Confirm Plans, Not Fill Fields</h2>\n<p>There is a larger shift here: what should an operations person actually be responsible for?</p>\n<p>In many admin workflows today, operations people carry too much mechanical work. The business goal is “run a May Day campaign”, but inside the admin system it becomes filling in a pile of fields, copying material URLs, and checking a pile of switches against a document.</p>\n<p>What people should really judge is not how every parameter is assembled.</p>\n<p>People should confirm the plan: who sees the campaign, when it starts, how much budget it uses, whether inventory is enough, whether it conflicts with another campaign, how it can be stopped, and which actions need approval.</p>\n<p>How many records that plan creates, which APIs it calls, and which systems it syncs to can be handled by an agent.</p>\n<p>A more reasonable flow would look like this:</p>\n<p>Operations states the goal. The agent generates a configuration plan.<br>\nThe CLI runs a dry-run and lists what will change.<br>\nThe human reviews impact, risk, and rollback.<br>\nAfter approval, execution happens.<br>\nAfter execution, the agent checks the result.</p>\n<p>This does not remove people from the process. It moves people away from being form-filling machines.</p>\n<p><img src=\"/assets/posts/2026/admin-cli-skill/02-dry-run-approval.jpg\" alt=\"A dry-run plan, risk review, and approval gate shown in one review surface\"></p>\n<p>Manual operations went through a similar transition. At first, someone logged into machines and typed commands. Later, the work moved into scripts, pipelines, approvals, and rollbacks. Humans remained, but they no longer had to memorize every tiny step.</p>\n<p>Operations admin systems will probably walk the same road.</p>\n<h2>Skill Is Where Experience Lives</h2>\n<p>CLI alone is not enough.</p>\n<p>CLI says what the platform can do. The difficult part of admin work is often not whether an API can be called, but whether it should be called in this situation.</p>\n<p>For example, creating a campaign may look simple as a command:</p>\n<pre><code class=\"language-bash\">ops campaign create --name &quot;May Day Campaign&quot; --city shanghai --start 2026-05-01 --end 2026-05-05\n</code></pre>\n<p>The hard parts are elsewhere.</p>\n<p>Can a new-user campaign overlap with a reactivation campaign? At what budget does approval become mandatory? Can a gray release city be expanded directly to nationwide traffic? Should material size be checked before a slot goes live? Should last week’s data be reviewed before changing a policy? Why does an old teammate always say “do not touch this field” even though the page allows it?</p>\n<p>That knowledge used to be scattered everywhere: PRDs, internal documents, group chats, page hints, code comments, and the kind of experience everyone knows but nobody quite writes down.</p>\n<p>That is what Skill should catch.</p>\n<p>By Skill, I do not mean renaming help documents. It should be closer to an operating procedure: when this kind of task appears, what should be checked first, which command should be run next, which parameters require human confirmation, which output must be verified, and how to recover after failure.</p>\n<p>Without Skill, an agent merely knows how to call APIs.</p>\n<p>That is dangerous.</p>\n<p>It may know how to create a campaign, but not that a certain channel should not be published after 10 p.m. It may know how to batch update, but not that an export should be made first. It may know how to change a risk-control threshold, but not that the threshold is connected to a live complaint path.</p>\n<p>What admin systems need to preserve is not only pages and APIs, but judgment that smells like real production trouble.</p>\n<h2>Technical Configuration Platforms Have the Same Problem</h2>\n<p>This is not only about operations admin systems.</p>\n<p>Technical configuration platforms have the same problem, and maybe a more urgent one.</p>\n<p><img src=\"/assets/posts/2026/admin-cli-skill/03-config-platform-network.jpg\" alt=\"Multiple technical configuration platforms connected through a shared command layer\"></p>\n<p>AB experiments, gateway routes, recommendation policies, risk-control rules, permission systems, monitoring alerts, CI/CD, task scheduling, and feature flags may look like tools for engineers, algorithms, operations, and data teams. In practice, they also involve a large amount of configuration work.</p>\n<p>Many of these platforms eventually end up in a familiar place: another admin UI.</p>\n<p>The page is not the problem. Having only the page is the problem.</p>\n<p>Technical users are usually not afraid of command lines. These platforms also tend to care more about APIs, permissions, and audit trails. Yet their capabilities are often wrapped inside page after page, and agents have to take a detour to use them.</p>\n<p>The more practical issue is that technical configuration often crosses systems.</p>\n<p>One release may involve changing a feature flag, updating a configuration center, refreshing a gateway, checking monitoring, attaching alerts, and preparing a rollback. Humans jump across several systems by experience. If an agent does not have a unified CLI and Skill, it has to guess. Guessing around production systems does not sound pleasant.</p>\n<p>So the claim can be widened a bit:</p>\n<p>Any platform that changes configuration, changes data, or triggers tasks through APIs should consider CLI + Skill.</p>\n<p>It does not have to happen all at once, but the direction is worth setting.</p>\n<h2>Choosing Tools</h2>\n<p>Implementation depends on the starting point.</p>\n<p>If the platform already has stable APIs, especially an OpenAPI contract, generating an SDK or CLI from that contract is the most direct path. <a href=\"https://www.speakeasy.com/docs/cli-generation/create-cli\">Speakeasy</a> and <a href=\"https://www.stainless.com/docs/cli/\">Stainless</a> both support generating CLIs from OpenAPI. <a href=\"https://openapi-generator.tech/\">OpenAPI Generator</a> can also generate client code, although a good business CLI still needs another layer of design on top.</p>\n<p>If there is no stable API, or if the target is existing software, desktop tooling, or a legacy system, <a href=\"https://github.com/HKUDS/CLI-Anything\">HKUDS/CLI-Anything</a> is worth looking at. It is closer to adding an agent-usable CLI harness to existing software than generating CRUD commands for backend APIs. Its emphasis on JSON output, tests, preview, and Skill matters more than the simple fact that commands can be generated.</p>\n<p>MCP is also worth connecting, but I would not make MCP the only entrance.</p>\n<p>MCP is a good way to connect tools to models. CLI is more like the platform’s own operating surface. CLI can be run by people, scripts, CI, and agents. Once capabilities are lowered into CLI, connecting MCP or any other agent platform becomes more natural.</p>\n<p>My current preferred order is:</p>\n<p>First clean up the API contract, then build the CLI, then write the Skill, and only then connect agent platforms.</p>\n<p>The order is not absolute, but it is better not to lock the capability into one model platform from the beginning. Admin systems usually outlive one generation of agent frameworks.</p>\n<h2>Do Not Get Lazy About Login</h2>\n<p>Once a CLI can change admin data, login and permissions cannot be treated casually.</p>\n<p>The easiest thing is to give users long-lived tokens and ask them to put those tokens in local config files. That feels convenient in the short term and painful in the long term. Token leaks, excessive permissions, employee departure, and audit attribution are all old problems.</p>\n<p>A more reliable approach is to treat the CLI as a real client.</p>\n<p>If the local environment can open a browser, use OAuth/OIDC <a href=\"https://oauth.net/2/native-apps/\">Authorization Code + PKCE</a>. In remote machines or browserless environments, consider <a href=\"https://datatracker.ietf.org/doc/html/rfc8628\">Device Authorization Grant</a>, often called device code flow: the CLI gives a short code, and the human confirms it in a browser.</p>\n<p>Tokens should go into the system keychain or credential manager when possible, not plain text in a project directory. High-risk commands should generate a plan instead of executing directly, or require an approval ticket, second confirmation, impact scope, and rollback method.</p>\n<p>Permissions cannot be lazy either.</p>\n<p>An operator who can create campaigns should not automatically be able to change risk-control policies. Someone who can gray release should not automatically be able to release nationwide. Someone who can rollback should not automatically be able to delete historical data. An agent should inherit the current human’s permissions, not run around with a super-operator account.</p>\n<p>If that boundary is not held, CLI + Skill becomes not an efficiency tool, but an incident accelerator.</p>\n<h2>Admin Pages Will Not Disappear</h2>\n<p>I do not think admin pages will disappear.</p>\n<p>In fact, many pages will become more important.</p>\n<p>Their center of gravity will move from “letting humans enter every field” toward “helping humans understand plans and risks.” Which objects will AI change? How many users are affected? Why this configuration? Where is approval now? Did the result after execution look normal? These all need good GUI.</p>\n<p>People should not be expected to take business responsibility from a pile of JSON. When someone needs to make the call, a page is still a good cognitive tool.</p>\n<p>But high-frequency, rule-bound, verifiable, rollbackable configuration actions should not stay trapped inside forms forever.</p>\n<p>In the past, building admin systems meant wrapping APIs into pages so that humans could operate systems.</p>\n<p>The next step may be turning operations into auditable, authorized, replayable capabilities, so that agents can execute and humans can judge.</p>\n<p>GUI continues to serve humans. CLI serves agents. Skill preserves experience. Approval and audit trails hold the risk.</p>\n<p>This sounds simple, but real implementation will certainly have traps. How should commands be designed? How should permissions be sliced? How should Skills be maintained? Which scenarios can execute automatically, and which must stop for a human? All of these have to be tried one by one.</p>\n<p>I happen to have an admin system of my own. I will probably use it for an experiment later.</p>\n<p>If it turns out that this idea only looked good on paper, that will still be worth writing down. At least from here, the road of endlessly adding more admin forms no longer feels sufficient. In the AI era, admin platforms should not only be clickable by humans. They should also be runnable by agents.</p>\n","date_published":"2026-05-24T00:00:00.000Z","tags":["AI","Admin Platform","CLI","Skill","Product"],"language":"en"},{"id":"https://www.lihuanyu.com/posts/2026/%E4%B8%AD%E5%90%8E%E5%8F%B0%E7%9A%84%E4%B8%8B%E4%B8%80%E7%AB%99%E6%98%AFCLI%E5%92%8CSkill/","url":"https://www.lihuanyu.com/posts/2026/%E4%B8%AD%E5%90%8E%E5%8F%B0%E7%9A%84%E4%B8%8B%E4%B8%80%E7%AB%99%E6%98%AFCLI%E5%92%8CSkill/","title":"中后台的下一站，是 CLI 和 Skill","summary":"从后台表单、配置平台和 AI agent 的使用方式出发，聊聊为什么中后台下一步不该只继续堆页面，而应该把能力沉淀成 CLI 和 Skill。","content_html":"<p>最近看后台表单看得有点烦。</p>\n<p><img src=\"/assets/posts/2026/admin-cli-skill/01-admin-console-cli.jpg\" alt=\"后台表单和 CLI 并存的操作界面\"></p>\n<p>不是某个页面特别难用，也不是哪个组件库又做出了奇怪的交互。烦的是另一件事：很多后台看起来模块很多，拆开以后其实都在重复同一种动作。</p>\n<p>打开列表，点新增，填名称、时间、范围、优先级、预算、素材、跳转地址，保存，提交审批，等生效。换一个业务模块，字段变了，流程没太变。再换一个技术配置平台，也差不多：改开关、调阈值、选环境、看 diff、发布、回滚。</p>\n<p>中后台当然不是只有表单。它有列表、图表、审批流、权限、导入导出、日志、数据看板。但真正高频的操作，最后经常还是回到那几个朴素动作：</p>\n<p>读一批数据，改一批配置，触发一次任务。</p>\n<p>说得再直白一点，很多中后台就是给接口套了一层人能操作的壳。</p>\n<p>这层壳过去非常有必要。不能让运营拿着 curl 改活动，不能让客服连数据库改状态，也不能让业务同学记几十个接口参数。页面把风险挡住，把字段摊开，把权限和校验加上，让人至少知道自己在干什么。</p>\n<p>但 AI agent 出现以后，这个前提有点变了。</p>\n<p>页面还是要有，但页面不该是平台能力唯一的入口。</p>\n<p><a href=\"/en/posts/2026/the-next-step-for-admin-platforms-is-cli-and-skill/\">English version: The Next Step for Admin Platforms Is CLI and Skill</a></p>\n<h2>后台是在保护人</h2>\n<p>我以前做后台时，对表单复杂这件事还算宽容。</p>\n<p>一个活动配置页有几十个字段，并不一定是产品经理作恶。活动本来就有时间、库存、渠道、人群、素材、预算、风控、灰度和兜底。页面如果不把这些东西摊开，人就更容易出事故。</p>\n<p>很多看起来啰嗦的后台，其实是在保护操作人。</p>\n<p>它把“这个字段不能为空”“这个城市不能投”“这个预算要审批”“这个资源位不能同时挂两个活动”写进页面和后端校验里。运营点得慢一点，总比线上活动配错强。</p>\n<p>问题是，这套设计默认了一件事：所有细节都要由人一步一步完成。</p>\n<p>在没有 AI 的时候，这个默认没什么毛病。人不填表，谁来填？人不点发布，谁来发布？但现在 agent 已经可以读文档、读接口、生成参数、跑命令、看结果，再让人确认。继续把所有能力都藏在页面后面，就有点浪费。</p>\n<p>更尴尬的是，让 agent 操作 GUI 往往很别扭。</p>\n<p>按钮在哪、弹窗有没有出来、表格滚动到哪一页、字段是不是被折叠了，这些对人来说是界面问题，对 agent 来说是噪声。它不需要一个精致的侧边栏，也不需要一个 8px 圆角的按钮。它更需要一个稳定的入口：</p>\n<pre><code class=\"language-bash\">ops campaign preview --id 123 --format json\nops campaign publish --id 123 --dry-run\nops campaign rollback --id 123 --reason &quot;库存配置错误&quot;\n</code></pre>\n<p>这类命令不好看，但它适合机器读，也适合人查。</p>\n<h2>CLI 比页面更适合 agent</h2>\n<p>开发工具领域已经把这件事演示过很多遍了。</p>\n<p><code>git status</code>、<code>pnpm test</code>、<code>kubectl get pods -o json</code>、<code>gh pr create</code> 这些命令没有什么产品包装，但 agent 用起来很顺。原因很简单：输入是文本，输出可控，失败有退出码，过程能记录，命令之间还能串起来。</p>\n<p>中后台如果只是给人用，GUI 是最自然的选择。</p>\n<p>但如果后台也要给 agent 用，CLI 就很难绕开。</p>\n<p>CLI 的好处不是“高级”，而是足够土。它可以被 shell 调，可以被 CI 调，可以被日志系统记录，可以在本地复现。出了问题，不需要回忆“我当时点了哪个按钮”，而是能看到哪条命令、哪个参数、哪个 request id、哪次审批。</p>\n<p>这对后台很重要。</p>\n<p>后台操作通常不是玩具。一个配置可能影响活动收入，一个开关可能影响用户路径，一个策略阈值可能影响风控。agent 可以执行，但执行过程必须留下痕迹。CLI 在这件事上比 GUI 自动化靠谱得多。</p>\n<p>所以我越来越觉得，中后台未来应该有两套入口。</p>\n<p>GUI 给人看，CLI 给 agent 跑。</p>\n<p>不是互相替代，而是分工。</p>\n<h2>运营确认方案，不是填字段</h2>\n<p>这里还有一个更大的变化：运营到底应该负责什么。</p>\n<p>现在很多后台流程里，运营承担了太多机械工作。业务目标是“做一个五一活动”，最后落到后台里，变成填写一堆字段，复制一堆素材地址，对照文档检查一堆开关。</p>\n<p>人真正应该判断的，其实不是每个参数怎么拼。</p>\n<p>人应该确认的是方案：活动给谁看，什么时候开始，预算有多少，库存够不够，会不会和别的活动冲突，出问题怎么停，哪些动作需要审批。</p>\n<p>至于这个方案最终要创建几条记录、调几个接口、同步几个系统，完全可以交给 agent 做。</p>\n<p>更合理的流程应该是这样：</p>\n<p>运营说目标，agent 生成配置方案；<br>\nCLI 跑 dry-run，列出会改什么；<br>\n人看影响范围、风险和回滚方式；<br>\n审批通过后再执行；<br>\n执行完继续检查结果。</p>\n<p>这不是把人踢出流程，而是把人从填表机器的位置上挪出来。</p>\n<p><img src=\"/assets/posts/2026/admin-cli-skill/02-dry-run-approval.jpg\" alt=\"dry-run 计划、风险和审批被放在同一个审查界面里\"></p>\n<p>以前手工运维也是这样。一个人登录机器敲命令，后来变成脚本、流水线、审批、回滚。人还在，只是人不再负责记住每一个细碎步骤。</p>\n<p>运营后台大概率也会走这条路。</p>\n<h2>Skill 才是经验</h2>\n<p>只有 CLI 还不够。</p>\n<p>CLI 只能说明“能做什么”。但后台麻烦的地方，常常不在能不能调接口，而在什么时候该不该调。</p>\n<p>比如创建活动这件事，命令可能很简单：</p>\n<pre><code class=\"language-bash\">ops campaign create --name &quot;五一活动&quot; --city shanghai --start 2026-05-01 --end 2026-05-05\n</code></pre>\n<p>真正难的是别的东西。</p>\n<p>新用户活动能不能和召回活动叠加？预算超过多少必须审批？灰度城市能不能直接扩到全国？资源位上线前要不要检查素材尺寸？改策略之前要不要先看上周数据？某个字段页面上能改，但老同事为什么总说“这个别碰”？</p>\n<p>这些东西过去散在很多地方：PRD、飞书文档、群聊、页面提示、代码注释，还有一些说不清楚但大家都知道的经验。</p>\n<p>这才是 Skill 应该接住的东西。</p>\n<p>我理解的 Skill，不是把帮助文档换个名字。它应该更像一份操作规程：遇到这类任务，先查什么，再跑什么命令，哪些参数要停下来问人，哪些输出必须核对，失败以后怎么恢复。</p>\n<p>没有 Skill，agent 只是会调用接口。</p>\n<p>这很危险。</p>\n<p>它可能知道怎么创建活动，却不知道晚上十点以后不能发某个渠道；知道怎么批量更新，却不知道更新前先导出；知道怎么改风控阈值，却不知道这个阈值背后挂着一条线上投诉链路。</p>\n<p>后台真正该沉淀的，不只是页面和接口，还有这些带着坑味的判断。</p>\n<h2>技术配置平台也一样</h2>\n<p>这件事不只发生在运营后台。</p>\n<p>技术配置平台也一样，而且可能更急。</p>\n<p><img src=\"/assets/posts/2026/admin-cli-skill/03-config-platform-network.jpg\" alt=\"多个技术配置平台通过统一命令层连接起来\"></p>\n<p>AB 实验、网关路由、推荐策略、风控规则、权限系统、监控告警、CI/CD、任务调度、feature flag，这些东西表面上是给研发、算法、运维、数据同学用的工具，本质上也有大量配置动作。</p>\n<p>很多平台最后都会走向一个熟悉结局：再做一个后台。</p>\n<p>页面不是错。错的是只有页面。</p>\n<p>技术平台的使用者本来就不怕命令行，平台也通常更重视 API、权限和审计。结果能力却被包在一个个页面里，agent 想接进来，只能绕路。</p>\n<p>更现实的问题是，技术配置往往跨系统。</p>\n<p>一次上线可能要改 feature flag，更新配置中心，刷网关，确认监控，挂告警，准备回滚。人靠经验在几个系统之间跳来跳去，agent 如果没有统一的 CLI 和 Skill，就只能靠猜。猜这种事，放在生产系统里听着就不吉利。</p>\n<p>所以这个判断可以放宽一点：</p>\n<p>凡是通过接口改配置、改数据、触发任务的平台，都应该考虑 CLI + Skill。</p>\n<p>不一定一步到位，但方向值得定下来。</p>\n<h2>工具怎么选</h2>\n<p>落地时要分清楚场景。</p>\n<p>如果平台本来就有稳定 API，尤其是已经有 OpenAPI，那么先从 API 契约生成 SDK 或 CLI 是最顺的。<a href=\"https://www.speakeasy.com/docs/cli-generation/create-cli\">Speakeasy</a> 和 <a href=\"https://www.stainless.com/docs/cli/\">Stainless</a> 都有从 OpenAPI 生成 CLI 的能力。<a href=\"https://openapi-generator.tech/\">OpenAPI Generator</a> 也能生成客户端代码，只是离一个好用的业务 CLI 还差一层设计。</p>\n<p>如果没有稳定 API，或者面对的是现成软件、桌面工具、遗留系统，那可以看看 <a href=\"https://github.com/HKUDS/CLI-Anything\">HKUDS/CLI-Anything</a>。它更像是在给已有软件补一层 agent 能用的 CLI harness，而不是专门给后台接口生成 CRUD 命令。它强调 JSON 输出、测试、预览和 Skill，这几个点比“能不能自动生成命令”更重要。</p>\n<p>MCP 也值得接，但我不想把 MCP 当成唯一入口。</p>\n<p>MCP 是给模型接工具的好方式，CLI 则更像平台自己的操作面。CLI 能被人跑，被脚本跑，被 CI 跑，也能被 agent 跑。能力沉到 CLI 以后，再接 MCP 或其他 agent 平台都比较自然。</p>\n<p>我现在更倾向的顺序是：</p>\n<p>先把 API 契约整理清楚，再做 CLI，然后写 Skill，最后再接 agent 平台。</p>\n<p>顺序不绝对，但最好别一上来就把能力绑死在某个模型平台里。后台系统的寿命通常比一代 agent 框架长。</p>\n<h2>登录别偷懒</h2>\n<p>CLI 一旦能改后台数据，登录和权限就不能偷懒。</p>\n<p>最省事的做法，是给用户发一个长期 token，让他放到本地配置文件里。这个方案短期舒服，长期难受。token 泄露、权限过大、离职回收、审计对人，哪个都不是小事。</p>\n<p>更靠谱的方式，是把 CLI 当成正式客户端。</p>\n<p>本地能打开浏览器，就走 OAuth/OIDC 的 <a href=\"https://oauth.net/2/native-apps/\">Authorization Code + PKCE</a>。远程机器或无浏览器环境，可以考虑 <a href=\"https://datatracker.ietf.org/doc/html/rfc8628\">Device Authorization Grant</a>，也就是 device code flow：CLI 给一个短码，人去浏览器里确认。</p>\n<p>token 尽量放系统 keychain 或凭据管理器，不要明文扔在项目目录。高风险命令只生成计划，不直接执行；或者必须带审批单、二次确认、影响范围和回滚方式。</p>\n<p>权限也不能图省事。</p>\n<p>运营能建活动，不代表能改风控。能灰度发布，不代表能全量上线。能回滚，不代表能删历史数据。agent 应该继承当前人的权限，而不是拿一个“超级运营”的万能账号到处跑。</p>\n<p>这个边界如果守不住，CLI + Skill 就不是效率工具，而是事故加速器。</p>\n<h2>后台不会消失</h2>\n<p>我并不觉得中后台页面会消失。</p>\n<p>相反，很多页面会变得更重要。</p>\n<p>只是页面的重心会从“让人输入所有字段”，慢慢转向“让人看懂方案和风险”。AI 准备改哪些对象，影响多少用户，为什么这么配，审批到哪一步，执行后结果是否正常，这些都需要好的 GUI。</p>\n<p>人不适合在一堆 JSON 里承担业务责任。真正要拍板的时候，页面仍然是很好的认知工具。</p>\n<p>但那些高频、规则明确、可验证、可回滚的配置动作，不应该永远困在表单里。</p>\n<p>以前做后台，是把接口包成页面，让人可以操作系统。</p>\n<p>下一步更重要的，可能是把操作变成可审计、可授权、可回放的能力，让 agent 可以执行，让人负责判断。</p>\n<p>GUI 继续服务人，CLI 服务 agent，Skill 沉淀经验，审批和审计兜住风险。</p>\n<p>这套东西听起来不复杂，真正落地时肯定会有坑。命令怎么设计，权限怎么切，Skill 怎么维护，哪些场景允许自动执行，哪些必须停下来问人，都要一件件试。</p>\n<p>我自己正好也有一个管理后台，后面应该会拿它做一轮实验。</p>\n<p>如果跑下来证明这事只是想得美，那也值得记一笔。至少现在看，中后台继续堆表单的路已经不太够用了。AI 时代的后台，不该只给人点，也该给 agent 跑。</p>\n","date_published":"2026-05-24T00:00:00.000Z","tags":["AI","中后台","CLI","Skill","产品"],"language":"zh"},{"id":"https://www.lihuanyu.com/en/posts/2026/gpt-image-2-prompt-patterns/","url":"https://www.lihuanyu.com/en/posts/2026/gpt-image-2-prompt-patterns/","title":"gpt-image-2 Prompt Library: Technical Diagrams, Covers, and Visuals","summary":"Tested gpt-image-2 prompts for technical diagrams, infographics, editorial covers, miniature scenes, logos, and knowledge cards, organized by use case.","content_html":"<p>I have been collecting image-generation prompts for a while, but the useful ones are rarely just long strings of adjectives.</p>\n<p>A good prompt is closer to a design brief. It tells the model what the image is for, what should be readable, which details matter, what must not be invented, and where the user is expected to replace the input.</p>\n<p>This English edition is not a literal translation of the Chinese series. I rewrote the prompts in English, changed the example subjects where it made sense, and regenerated the images with <code>gpt-image-2</code>. The Chinese examples still matter, but English prompts deserve their own tests.</p>\n<p>The templates in this series are adapted from public prompt ideas shared by @xiaoxiaodong01 and @MrLarus, with thanks.</p>\n<p><a href=\"/posts/gpt-image-2-prompt-gallery/\">Chinese version of this article</a></p>\n<h2>Choose a Prompt Set</h2>\n<div class=\"table-scroll\"><table>\n<thead>\n<tr>\n<th>What you need</th>\n<th>Start here</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Architecture overview or mechanism explainer</td>\n<td><a href=\"/en/posts/2026/gpt-image-2-technical-diagram-prompts/\">Technical diagrams and infographics</a></td>\n</tr>\n<tr>\n<td>A visual reconstruction of a Mermaid flow</td>\n<td><a href=\"/en/posts/2026/gpt-image-2-technical-diagram-prompts/\">Technical diagrams and infographics</a></td>\n</tr>\n<tr>\n<td>Editorial cover or report visual</td>\n<td><a href=\"/en/posts/2026/gpt-image-2-commercial-cover-prompts/\">Commercial covers and miniature visuals</a></td>\n</tr>\n<tr>\n<td>Miniature product scene, concept logo, or knowledge card</td>\n<td><a href=\"/en/posts/2026/gpt-image-2-commercial-cover-prompts/\">Commercial covers and miniature visuals</a></td>\n</tr>\n</tbody>\n</table>\n</div><p>Each article includes the tested input, the generated result, the reusable prompt, and the failure modes that matter for that format. This page stays short so it can remain the stable entry point as the library grows.</p>\n<h2>What I Look For</h2>\n<p>I do not judge a prompt by length. Long prompts are easy to write. Reusable prompts are harder.</p>\n<p>The ones worth saving usually do a few things well:</p>\n<ul>\n<li>They know their use case: cover image, technical diagram, report visual, poster, wallpaper, or brand direction.</li>\n<li>They control information density instead of asking the model to include everything.</li>\n<li>They give the image a reading path: title, subject, supporting labels, negative space, and a closing point.</li>\n<li>They name failure modes: garbled text, fake data, crowded layouts, cheap templates, wrong components.</li>\n<li>They leave a clear input slot so the pattern can be reused.</li>\n</ul>\n<p>The pattern is simple: a good prompt does not merely describe a style. It assigns work.</p>\n<h2>Series Index</h2>\n<h3>Technical Diagrams And Infographics</h3>\n<p>For architecture overviews, mechanism explainers, workflow diagrams, Mermaid reconstruction, and technical presentation visuals.</p>\n<p><img src=\"/assets/posts/2026/gpt-image2-prompts-en/01-agent-architecture-handdrawn.jpg\" alt=\"Hand-drawn knowledge diagram for an agent runtime\"></p>\n<p>Read the full set: <a href=\"/en/posts/2026/gpt-image-2-technical-diagram-prompts/\">gpt-image-2 Prompt Patterns: Technical Diagrams And Infographics</a></p>\n<h3>Commercial Covers And Miniature Visuals</h3>\n<p>For editorial information visuals, report covers, technology posters, miniature product scenes, brand concepts, and knowledge cards.</p>\n<p><img src=\"/assets/posts/2026/gpt-image2-prompts-en/06-lavender-knowledge-card.jpg\" alt=\"Lavender botanical knowledge card\"></p>\n<p>Read the full set: <a href=\"/en/posts/2026/gpt-image-2-commercial-cover-prompts/\">gpt-image-2 Prompt Patterns: Commercial Covers And Miniature Visuals</a></p>\n<h2>A Practical Note</h2>\n<p>Image prompts age differently from normal code.</p>\n<p>Model names, supported sizes, pricing, and platform entry points will change. The durable part is the prompt structure: what the image is supposed to do, what information is allowed, how text is handled, and how the final result will be used.</p>\n<p>Text inside generated images still needs manual checking. Large titles and short labels are often usable. Dense annotations, names, dates, numbers, and specialized terms should be verified before publishing.</p>\n","date_published":"2026-05-23T00:00:00.000Z","date_modified":"2026-07-31T00:00:00.000Z","tags":["AI","Prompting","Image Generation","gpt-image-2","Design"],"language":"en"},{"id":"https://www.lihuanyu.com/en/posts/2026/gpt-image-2-commercial-cover-prompts/","url":"https://www.lihuanyu.com/en/posts/2026/gpt-image-2-commercial-cover-prompts/","title":"gpt-image-2 Prompts for Commercial Covers and Miniature Visuals","summary":"Tested gpt-image-2 prompts for editorial covers, miniature product scenes, technology posters, concept logos, and botanical knowledge cards.","content_html":"<p>There is a thin line between useful content imagery and empty commercial polish.</p>\n<p>If the image looks too much like an advertisement, it feels oily. If it behaves like a plain data chart, it may not travel. The prompts here are for the middle ground: report covers, editorial visuals, miniature scenes, brand directions, and knowledge cards.</p>\n<p>These examples use English prompts and English subjects. They are adapted from public prompt ideas shared by @xiaoxiaodong01 and @MrLarus, then rewritten and tested again with <code>gpt-image-2</code>.</p>\n<p><a href=\"/posts/gpt-image-2-commercial-cover-prompts/\">Chinese version of this article</a></p>\n<p><a href=\"/en/posts/2026/gpt-image-2-prompt-patterns/\">Series index</a></p>\n<h2>Choose a Pattern</h2>\n<div class=\"table-scroll\"><table>\n<thead>\n<tr>\n<th>Output</th>\n<th>Best starting point</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Report cover or editorial information visual</td>\n<td>Editorial information visual</td>\n</tr>\n<tr>\n<td>Technology poster or product-themed social image</td>\n<td>Electronics miniature typography poster</td>\n</tr>\n<tr>\n<td>Early brand direction or abstract product mark</td>\n<td>Conceptual narrative minimal logo</td>\n</tr>\n<tr>\n<td>Calm educational card with a single subject</td>\n<td>Quiet botanical knowledge card</td>\n</tr>\n</tbody>\n</table>\n</div><p>These prompts solve different communication problems. Pick the output format first, then replace the subject and source material inside that pattern instead of combining every visual style into one request.</p>\n<h2>Case 1: Editorial Information Visual</h2>\n<p>Some topics do not want to become a flowchart.</p>\n<p>I used <em>Moby-Dick</em> as the test subject. It is not a report, and it does not have a system architecture. But it has structure: Ahab, Ishmael, the Pequod, the white whale, obsession, fate, the ocean, and narrative drift.</p>\n<p>This prompt turns that structure into a micro-landscape. The result is not a chart. It is a readable editorial scene with platforms, labels, a ship, a whale, and a final line. The labels are mostly readable, but any generated literary notes should still be checked before publishing.</p>\n<p><img src=\"/assets/posts/2026/gpt-image2-prompts-en/03-moby-dick-info-visual.jpg\" alt=\"Editorial information visual for Moby-Dick\"></p>\n<p>Generation notes:</p>\n<ul>\n<li>Model: <code>gpt-image-2</code></li>\n<li>Size: <code>1536x1024</code></li>\n<li>Output: JPEG, about 211 KB after compression</li>\n<li>Test input: <em>Moby-Dick</em></li>\n<li>Best for: presentation covers, report illustrations, knowledge covers, editorial visuals</li>\n</ul>\n<p>Reusable prompt:</p>\n<pre><code class=\"language-text\">Create an advanced editorial information visual image.\n\nIt should not look like a normal chart or a template infographic. Transform information, ideas, text, and emotion into a bright, restrained micro-landscape with spatial depth.\n\nBefore visual creation:\n- If the input includes data, screenshots, tables, reports, images, or explicit source material, use those materials first.\n- If the input only gives a topic, year, industry, issue, or direction without enough data, use only stable and verifiable knowledge.\n- Do not invent current market data, future trends, prices, policies, rankings, people, organizations, brands, or news.\n- The image may be poetic, but the information must be honest.\n\nVisual structure:\n- Turn the information into visible geometric objects: cubes, thin columns, slices, platforms, containers, boundaries, suspended pieces, folded structures, small markers, and spatial layers.\n- Important information may have clearer volume, higher placement, and more stable structure.\n- Secondary information should feel lighter, lower, and closer to the background.\n- Hidden or subtle information can live in edges, gaps, folds, shadows, labels, and fine lines.\n- Do not make the image look like auto-generated software charts.\n\nComposition:\n- Start from abundant negative space.\n- The subject may be centered or slightly off-center.\n- Use a stable but light visual support: a pale platform, thin layer, transparent boundary, floating base, or abstract object.\n- Other information elements may surround it, pass through it, be lightly occluded by it, or extend from it.\n- Keep a hint of perspective and volume, but avoid realistic 3D rendering.\n- The final texture should sit between flat illustration, print design, low-poly paper forms, and editorial design.\n\nColor:\n- Do not use a fixed palette mechanically.\n- Generate a color relationship from the topic's mood, information density, use case, and amount of whitespace.\n- Keep the image bright, clean, airy, and paper-like.\n- Use value, saturation, temperature, transparency, area, and spatial distance to separate hierarchy.\n- Dark colors should only appear in small text, fine lines, edges, ticks, local shadows, or visual pauses.\n- Avoid muddy colors, heavy darkness, excessive sweetness, commercial gloss, neon color, and template-like palettes.\n\nText:\n- Text is part of the image structure, not a manual.\n- Titles, short phrases, numbers, labels, and annotations should be embedded in objects, edges, whitespace, and reading flow.\n- Important text may occupy a clean area of negative space.\n- Secondary text may sit on columns, folds, side edges, pale shadows, or small components.\n- Text should appear only where it helps.\n- Avoid filling the image with text.\n\nHuman scale:\n- Add a few tiny people if useful.\n- They may stand, observe, measure, carry, look upward, pass by, or pause.\n- They exist to give scale and make abstract information feel human.\n- Keep them simple and quiet.\n\nFinal use:\nThe image should work as a presentation cover, report illustration, information graphic, social cover, brand visual, business card background, or knowledge content visual. It should carry information while still feeling like an image worth looking at.\n\nUser input:\n{topic, core expression, source text or data, reference image if any, intended use}\n</code></pre>\n<h2>Case 2: Electronics Miniature Typography Poster</h2>\n<p>Short lines often become cheap light-and-shadow posters. The miniature electronics constraint helps.</p>\n<p>I used the sentence “The sun waits at the edge of night.” The generated image turns the sentence into a threshold: electrical modules, a glowing door, a small figure, and a large readable title. That is exactly the point of this pattern. The text is not decoration; it is one of the main subjects.</p>\n<p><img src=\"/assets/posts/2026/gpt-image2-prompts-en/04-electronics-dawn-poster.jpg\" alt=\"Electronics miniature poster with a dawn threshold\"></p>\n<p>Generation notes:</p>\n<ul>\n<li>Model: <code>gpt-image-2</code></li>\n<li>Size: <code>1536x1024</code></li>\n<li>Output: JPEG, about 157 KB after compression</li>\n<li>Test input: “The sun waits at the edge of night.”</li>\n<li>Best for: social covers, presentation openings, technology brand visuals, exhibition posters</li>\n</ul>\n<p>Reusable prompt:</p>\n<pre><code class=\"language-text\">Translate the user input into an electronics miniature typography poster.\n\nFirst understand the real communication task:\nIs this a cover, infographic, presentation opening slide, brand visual, product explanation, exhibition poster, social image, or business card background?\n\nThe image must be both beautiful and readable. If the user provides text, that text is one of the main subjects, not a small decoration.\n\nText strategy:\n- Base all visible text on the user input.\n- Do not add unrelated slogans, fake data, hollow industry words, or random English.\n- If the user provides a complete line, optimize hierarchy and layout without changing its meaning.\n- If the user provides keywords, add only a small number of relevant short labels.\n- If the user provides no text, generate only restrained and accurate thematic text.\n- Main title, subtitle, labels, and notes must have a clear hierarchy.\n- Important text must not be too small, hidden in a corner, or swallowed by the scene.\n\nText integration:\n- Text may appear on building faces, boards, miniature light boxes, signs, floor wayfinding, transparent information layers, product labels, margins, or callout annotations.\n- It may also become a strong main title.\n- Typography should fit the theme: technical, warm, industrial, editorial, playful, premium, exhibition-like, or social-cover impact.\n- Letter spacing, alignment, weight, and whitespace should feel professionally designed.\n\nMiniature scene:\n- Build the space from recognizable electronic components: sockets, power strips, switches, circuit breakers, cables, terminals, relays, power modules, plugs, and wiring structures.\n- Components may become a building, street, station, workshop, booth, tower, track, road, bridge, or running system.\n- Keep real material and recognizable structure.\n- Do not deform components into strange unrecognizable objects.\n\nComposition:\n- Include one visual core.\n- Include one clear text reading area.\n- Include one miniature human narrative area.\n- Include one quieter auxiliary information area.\n- Use center, split, layered, diagonal, title-in-whitespace, or modular composition according to the user's intention.\n- Keep breathing room. Do not fill the entire image evenly.\n\nMiniature people:\n- Use small people as narrative clues.\n- They may repair, build, carry, present, collaborate, commute, observe, queue, or work around the core device.\n- Their actions should be small and precise.\n- Plants, tracks, copper wires, roads, steps, tools, and signs may guide the eye and soften the industrial feeling.\n\nVisual metaphor:\nInfer the metaphor from the theme: growth, connection, safety, energy, efficiency, collaboration, brand, service, manufacturing, or supply chain.\nExpress it through spatial relationships, human behavior, component structure, and text hierarchy. Do not explain it directly.\n\nStyle:\n- Miniature model photography\n- Product advertising photography\n- Shallow depth of field\n- 3/4 overhead view\n- Soft studio lighting\n- Precise plastic and metal texture\n- Clean color palette that supports readability\n\nAvoid:\nreal brand logos, garbled text, unrelated copy, fake data, empty slogans, tiny keywords, hidden text, pure scenery without information hierarchy, crowded layout, dirty circuit boards, cyberpunk neon, wasteland mood, cheap toy feeling, unrecognizable components, and oversized miniature people.\n\nUser input:\n{phrase, topic, or short copy}\n</code></pre>\n<h2>Case 3: Conceptual Narrative Minimal Logo</h2>\n<p>Logo prompts are easy to overtrust.</p>\n<p>Image models are good at making pictures that look like logos. A real mark still needs vector work, small-size checks, and brand-system thinking. This prompt is more useful as a concept generator: it asks for metaphor, negative space, typography, and a quiet relationship between symbol and name.</p>\n<p>I tested it with InkIsle. The result gives a clear direction: ink, island, page, Markdown, route, and a small red confirmation point. It is not a final logo file, but it is a useful brand concept draft.</p>\n<p><img src=\"/assets/posts/2026/gpt-image2-prompts-en/05-inkisle-concept-logo-en.jpg\" alt=\"Conceptual narrative minimal logo concept for InkIsle\"></p>\n<p>Generation notes:</p>\n<ul>\n<li>Model: <code>gpt-image-2</code></li>\n<li>Size: <code>1024x1024</code></li>\n<li>Output: JPEG, about 48 KB after compression</li>\n<li>Test input: InkIsle</li>\n<li>Best for: brand concept drafts, independent studio identity, creative publishing tools, art project marks</li>\n</ul>\n<p>Reusable prompt:</p>\n<pre><code class=\"language-text\">Design a high-completion conceptual narrative minimal logo.\n\nUser input:\n- Brand or project name: {brand name}\n- Subtitle or product line: {subtitle}\n- Type / industry: {studio, software, publication, art project, fashion label, music brand, architecture space, cultural project, independent brand, etc.}\n- Positioning: {brand positioning}\n- Core concept: {dream, control, distance, solitude, exploration, future, imagination, dialogue, structure, freedom, narrative, spirituality, etc.}\n- Metaphor: {astronaut, rabbit ears, silhouette, puppet, hand, box, planet, door, eye, thread, moon, installation, geometric structure, figure-space relationship, etc.}\n- Mood: {restrained, calm, experimental, mysterious, poetic, rational, futuristic, independent, etc.}\n- Main colors: {black, white, gray, deep brown, ink black, etc.}\n- Accent color: {small red, dark red, or another restrained accent}\n- Aspect ratio: 1:1\n\nGoal:\nThis is not a normal corporate logo and not a simple icon. It should be a small visual work with concept, narrative tension, and minimal character.\n\nThe result should feel:\n- minimal but not empty\n- spacious but not weak\n- simple but meaningful\n- restrained but designed\n- closer to an independent studio or art project identity than a commercial badge\n\nDesign essence:\nDo not directly explain what the brand does. Express the brand's spirit indirectly through metaphor, scene relationship, and visual poetry.\n\nPrioritize:\n- whether the mark has a concept\n- whether symbol and text form a narrative relationship\n- whether negative space strengthens the work\n- whether the whole image feels like a small art-directed identity system\n\nSymbol:\n- Design a small symbolic mark, scene, or installation from the core concept and metaphor.\n- It may use black-and-white silhouettes, fine line structure, geometric frames, tiny red lines, or selective negative space.\n- It must be simple, but not generic.\n- It should have viewing value, not only functional icon value.\n\nTypography:\n- The brand name must be clearly readable.\n- English may use restrained thin, modern, uppercase or editorial typography.\n- Chinese or bilingual text may be used only if the input requires it.\n- Text can act as annotation, counterpoint, title, support, or structural element.\n- Avoid decorative fonts and crowded text.\n- Symbol and text must feel intentionally related.\n\nComposition:\n- Lots of negative space.\n- The main subject may be small.\n- Centered, off-center, split, opposing, suspended, or aligned layouts are all allowed if balanced.\n- Fine lines, tiny markers, numbers, or subtitles may appear sparingly.\n- The image should feel like a brand identity proposal, not a poster or packaging front.\n\nColor:\n- Use black, white, gray, deep brown, or ink-like tones.\n- Use only a tiny accent color, such as red or dark red, for a route, connection, warning point, or relationship marker.\n- Avoid high saturation, large gradients, noisy colors, and commercial shine.\n\nAvoid:\ntraditional heavy corporate badges, cartoon logos, generic geometric marks, crowded layouts, glossy effects, random symbols, fake slogans, and cheap logo-template aesthetics.\n\nFinal output:\nA polished conceptual narrative minimal logo with metaphor, negative space, restrained typography, and independent studio character.\n</code></pre>\n<h2>Case 4: Quiet Botanical Knowledge Card</h2>\n<p>This pattern is for light knowledge visuals.</p>\n<p>It is not a dense botanical encyclopedia page and not a dashboard. The image should feel like a plant placed on paper, with a few pieces of information allowed to stop around it.</p>\n<p>I tested it with Lavender. The image keeps six knowledge points, gives the plant a large central presence, and uses quiet labels instead of heavy information blocks. This kind of image works well for a presentation cover, a knowledge card, or a visual index page.</p>\n<p><img src=\"/assets/posts/2026/gpt-image2-prompts-en/06-lavender-knowledge-card.jpg\" alt=\"Lavender botanical knowledge card\"></p>\n<p>Generation notes:</p>\n<ul>\n<li>Model: <code>gpt-image-2</code></li>\n<li>Size: <code>2048x1024</code></li>\n<li>Output: JPEG, about 256 KB after compression</li>\n<li>Test input: Lavender</li>\n<li>Best for: botanical cards, presentation covers, knowledge cards, social covers, brand visuals</li>\n</ul>\n<p>Reusable prompt:</p>\n<pre><code class=\"language-text\">Create a restrained, translucent, negative-space knowledge card.\n\nThe image should work for a presentation, infographic, social cover, knowledge card, business card, or brand visual. It should not chase decoration. It should feel like a quiet generated object: information is gently lifted, space breathes, and visual elements avoid and attract each other at the same time.\n\nColor:\n- Use low-saturation, pale, misty, paper-like colors.\n- Background should be warm off-white, pale warm gray, mist gray, or light beige-gray. Do not use dead pure white.\n- Add one soft translucent pale mint, ice blue-green, pale cyan, or mist-blue shape as a spatial support.\n- This shape should not feel like a hard standard geometric block; it should feel diluted by light.\n- Text color should be smoky gray, deep brown-gray, or gray-black.\n- Accent colors may come from nature, but they must be light, few, and precise.\n\nComposition:\n- Use asymmetrical negative space unless the user asks for centered composition.\n- The main subject should not fill the canvas mechanically.\n- It may grow diagonally, breathe vertically, or shift slightly from center.\n- Use a natural branch, soft line, plant, delicate curve, mist shape, paper texture, pale shadow, or abstract data trace as the visual line.\n- The elements should feel naturally grown rather than mechanically placed.\n\nTypography:\n- Typography is the skeleton of the image.\n- Titles may be vertical, segmented, or placed with generous letter spacing.\n- English notes should use light sans-serif or restrained serif typography, small size, and wide tracking.\n- Body information should be broken into short lines, poetic fragments, data labels, or small information nodes.\n- Important numbers may be enlarged, but should remain quiet and elegant.\n- Text and image should support each other.\n\nIf the input is knowledge, data, opinions, or report content:\n- Turn it into a lightweight information graphic.\n- Use a small number of lines, pale color blocks, labels, numeric hierarchy, vertical rhythm, and whitespace grouping.\n- Do not create a traditional table, dense flowchart, or mechanical dashboard.\n- Every group of text should have its own pause and position.\n\nLight and texture:\n- Soft natural diffused light\n- Low contrast\n- Slight depth of field\n- Subtle paper texture\n- Mild air particles\n- Semi-transparent layers\n- Restrained edges\n\nAvoid:\nneon, cyberpunk, high-saturation gradients, heavy 3D, cartoon style, hard tech wireframes, standard icon templates, commercial stock collage, crowded text, and aggressive contrast.\n\nUser variables:\n- Topic: {topic}\n- Purpose: {use case}\n- Knowledge points: {number and requirements}\n- Composition: {centered, large subject, diagonal growth, etc.}\n- Aspect ratio: {ratio}\n</code></pre>\n","date_published":"2026-05-23T00:00:00.000Z","date_modified":"2026-07-31T00:00:00.000Z","tags":["AI","Prompting","Image Generation","gpt-image-2","Design"],"language":"en"},{"id":"https://www.lihuanyu.com/en/posts/2026/gpt-image-2-technical-diagram-prompts/","url":"https://www.lihuanyu.com/en/posts/2026/gpt-image-2-technical-diagram-prompts/","title":"gpt-image-2 Prompts for Technical Diagrams and Infographics","summary":"Tested gpt-image-2 prompts for turning architecture notes, technical mechanisms, and Mermaid flows into readable technical diagrams and infographics.","content_html":"<p>Technical images fail in two opposite ways.</p>\n<p>One version includes everything and helps no one. The other looks polished but quietly breaks the mechanism it is supposed to explain.</p>\n<p>The prompts here try to solve that by forcing the model to compress first, then design. The main path, node roles, arrows, labels, and bottom-line conclusion matter more than decoration.</p>\n<p>These examples use English prompts and English inputs. They are adapted from public prompt ideas shared by @xiaoxiaodong01 and @MrLarus, then rewritten and tested again with <code>gpt-image-2</code>.</p>\n<p><a href=\"/posts/gpt-image-2-technical-diagram-prompts/\">Chinese version of this article</a></p>\n<p><a href=\"/en/posts/2026/gpt-image-2-prompt-patterns/\">Series index</a></p>\n<h2>Choose a Pattern</h2>\n<div class=\"table-scroll\"><table>\n<thead>\n<tr>\n<th>Input</th>\n<th>Best starting point</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Architecture notes, modules, and system boundaries</td>\n<td>Hand-drawn knowledge diagram</td>\n</tr>\n<tr>\n<td>A mechanism that needs a clear reading path</td>\n<td>Hand-drawn knowledge diagram</td>\n</tr>\n<tr>\n<td>Existing Mermaid source with nodes and arrows</td>\n<td>Technical infographic reconstruction</td>\n</tr>\n<tr>\n<td>A workflow that must preserve sequence and relationships</td>\n<td>Technical infographic reconstruction</td>\n</tr>\n</tbody>\n</table>\n</div><p>The first pattern compresses a technical subject into a readable knowledge card. The second preserves an existing graph while replacing the plain flowchart treatment with stronger hierarchy and visual grouping.</p>\n<h2>Case 1: Hand-Drawn Knowledge Diagram</h2>\n<p>Developers often need diagrams that explain a mechanism without becoming a cold flowchart.</p>\n<p>I tested this prompt with an “Atlas Agent Runtime Overview”. The useful part is that the prompt asks the model to preserve the technical path while turning the content into a readable hand-drawn explainer: title, judgment, modules, arrows, observability layer, and a bottom flow summary.</p>\n<p>The result keeps the main path visible and the English text is mostly readable. For production use, I would still keep the input under six modules. Beyond that, the model starts trading readability for completeness.</p>\n<p><img src=\"/assets/posts/2026/gpt-image2-prompts-en/01-agent-architecture-handdrawn.jpg\" alt=\"Hand-drawn knowledge diagram for an Atlas Agent Runtime\"></p>\n<p>Generation notes:</p>\n<ul>\n<li>Model: <code>gpt-image-2</code></li>\n<li>Size: <code>1536x1024</code></li>\n<li>Output: JPEG, about 297 KB after compression</li>\n<li>Test input: Atlas Agent Runtime Overview</li>\n<li>Best for: architecture summaries, technical explainers, retrospectives, internal sharing visuals, knowledge cards</li>\n</ul>\n<p>Test input:</p>\n<pre><code class=\"language-text\">Atlas Agent Runtime Overview\n\nCore judgment:\nAn agent is not a chat box; it is an observable execution loop that turns intent into plan, tool calls, memory updates, and deliverable output.\n\nMain path:\nUser Request -&gt; Orchestrator -&gt; Planner -&gt; Tool Router -&gt; Tools and APIs -&gt; Memory / State -&gt; Response Builder -&gt; User Review\n\nKey modules:\n1. Orchestrator receives requests, preserves context, and decides the next module.\n2. Planner breaks goals into steps, dependencies, risks, and missing information.\n3. Tool Router selects code, search, documents, database, browser, and internal APIs.\n4. Memory / State stores user preferences, task state, intermediate artifacts, and reusable context.\n5. Guardrails / Observability records tool_call, errors, latency, permission boundaries, and rollback points.\n6. Response Builder turns execution results into an answer, code change, or next-step proposal.\n\nBottom Line:\nA maintainable agent makes every decision, every tool call, and every state change traceable.\n</code></pre>\n<p>Reusable prompt:</p>\n<pre><code class=\"language-text\">Create a highly readable hand-drawn knowledge diagram from the content I provide.\n\nThe style should feel like a carefully organized creative notebook, a whiteboard walkthrough, and a consulting-style explainer. It should not look like a cold template or a generic flowchart.\n\nGoal:\nGenerate a diagram suitable for sharing, presentation, and reuse. The viewer should first understand the core judgment, then follow the modules, and finally remember one bottom-line conclusion.\n\nLanguage:\nAll visible text should follow the language of the input. Do not mix languages unless the input contains product names, protocol names, code tokens, paths, or numeric metrics.\n\nCanvas:\n- Aspect ratio: {16:9 / 5:4 / 4:3 / 21:9}\n- High-resolution output\n- Warm off-white or light warm-gray background\n- Subtle paper texture and enough breathing room\n- Keep all text readable; do not squeeze content into tiny labels\n\nInformation design:\n- Do not copy the source text word for word. Compress first, then draw.\n- Structure the image as:\n  1. Top: strong title plus one-sentence core judgment\n  2. Middle: 3 to 6 main modules arranged by flow, comparison, phase, or cause and effect\n  3. Inside each module: up to 3 to 5 short bullets\n  4. Bottom: one Flow Summary, Decision Summary, or Bottom Line\n  5. If the input is dense, keep only the 8 to 10 most important judgments\n\nReadability:\n- The title must be the largest and clearest element.\n- Module titles should feel ordered.\n- Body text must be short.\n- Each module should avoid more than six body lines.\n- Avoid dense tables and tiny technical paragraphs.\n- Do not sacrifice readability for completeness.\n\nVisual style:\n- Use dark ink or black hand-drawn lines to build the reading structure.\n- Use rounded sections, fine frames, light shadows, numbering, arrows, labels, and small icons.\n- Lines may have slight hand-drawn irregularity, but alignment, margins, and grouping must stay stable.\n- Icons are secondary wayfinding elements, not the main visual hierarchy.\n\nColor:\n- Warm paper background plus dark ink lines.\n- Use restrained marker colors such as low-saturation teal, sage, muted lavender, soft orange, and pale blue.\n- Avoid neon colors, strong gradients, heavy commercial glow, and one-color monotony.\n- Colored areas should occupy a small to medium part of the image.\n\nAccuracy:\n- Preserve the technical chain, component names, arrow direction, protocols, ports, data flow, and judgments in the input.\n- Do not invent components that were not provided.\n- Do not reverse actions. For example, &quot;read logs&quot; must not become &quot;generate logs&quot;.\n- If space is limited, preserve the main path, key distinctions, and final judgment; remove secondary explanations.\n\nContent:\n{paste your content here}\n</code></pre>\n<h2>Case 2: Reconstructing A Mermaid Flow As A Technical Infographic</h2>\n<p>Mermaid is great inside documentation. It is not always good enough for a presentation slide or article cover.</p>\n<p>I tested this with an image asset pipeline: write <code>prompts.jsonl</code>, generate raw images, compress them, publish assets, upload to a CDN, and reference them from Markdown. The original Mermaid flow is already clear, but visually it still feels like an engineering sketch.</p>\n<p>This prompt does not ask for “a prettier Mermaid chart”. It asks the model to extract the mechanism, assign node roles, and redraw the flow as a technical editorial infographic.</p>\n<p><img src=\"/assets/posts/2026/gpt-image2-prompts-en/02-image-pipeline-infographic.jpg\" alt=\"Technical infographic reconstruction of an image asset pipeline\"></p>\n<p>Generation notes:</p>\n<ul>\n<li>Model: <code>gpt-image-2</code></li>\n<li>Size: <code>1536x1024</code></li>\n<li>Output: JPEG, about 128 KB after compression</li>\n<li>Test input: a Mermaid flow for an image asset pipeline</li>\n<li>Best for: technical documents, architecture walkthroughs, workflow explanations, presentation slides</li>\n</ul>\n<p>Test input:</p>\n<pre><code class=\"language-mermaid\">flowchart LR\n  A[Write prompts.jsonl] --&gt; B[imgasset generate]\n  B --&gt; C[raw images]\n  C --&gt; D[TinyPNG compression]\n  D --&gt; E[publish assets]\n  E --&gt; F[CDN upload]\n  F --&gt; G[Markdown reference]\n</code></pre>\n<p>Reusable prompt:</p>\n<pre><code class=\"language-text\">You are a senior information architect and technical infographic designer.\n\nTask:\nRedesign the provided Mermaid / C4 / Flowchart / Sequence / State / ER / Timeline source, or a rendered diagram image, into a high-fidelity professional technical infographic.\n\nThe goal is not to beautify the original diagram, not to replicate Mermaid, and not to reskin a normal flowchart. First understand the semantic structure, then recompile it into a new information architecture.\n\nOutput:\nGenerate only the final infographic. Do not output analysis, Markdown, source code, or design notes.\n\nInput truth rules:\n1. If the input is source code, the source is the semantic truth. Ignore original Mermaid layout, colors, class definitions, and node styles.\n2. If the input is an image, use it only for semantic extraction. Do not copy the original layout, palette, node shapes, or arrow paths.\n3. If the source image is blurry, keep only confirmed information. Do not invent business logic.\n\nSemantic extraction:\nExtract entities, groups, actors, relationships, branches, merges, loops, gates, tools, stores, schemas, states, outputs, triggers, annotations, dependencies, and observations.\n\nRole assignment:\nAssign a role to each entity:\ninput, output, controller, orchestrator, processor, resolver, decision, gate, tool, storage, observer, actor, artifact, annotation, boundary, event, state, terminal, or reference.\n\nPrimary mechanism:\nIdentify one primary mechanism and make it visible within three seconds.\nPossible mechanisms include:\npipeline, orchestration, resolver pipeline, gating, handoff, layered system, lifecycle, hub-and-spoke, dependency network, sequence interaction, state transition, artifact-centered flow, split decision tree, release flow, or deployment flow.\n\nPrimary path:\nExtract the main path:\ninput -&gt; controller / processor / decision / resolver -&gt; output.\nMake this the visual spine. Auxiliary relationships should become branches, loops, lookup lines, observation lines, dependency lines, or side annotations.\n\nTemplate selection:\n- Linear primary path with 7 or fewer major steps: Core Flow Spine\n- Central controller with 3 or more branches: Orchestrator Hub\n- Runtime / gateway / engine / tools / storage: Layered Blueprint\n- Validation or conditional branching: Split Gate\n- Multi-actor interaction: Swim Relay\n- User action plus internal state toggle: Interaction State Panel\n- Artifact or version resolves execution: Artifact Anchor Resolver\n- Entity relationship: Relational Data Grid\n- Dense dependency graph: Clustered Zones\n- Lifecycle or state transition: Lifecycle Ring\n\nStyle:\nUse a premium technical editorial style by default:\n- warm off-white or soft technical canvas\n- restrained surfaces\n- modern clean sans-serif typography\n- low-saturation navy, teal, amber, and gray accents\n- subtle paper grain or micro-grid if useful\n- no heavy shadow, glow, glassmorphism, 3D, rainbow palette, or loud gradient\n\nInformation compression:\n- Merge similar nodes when there are more than four.\n- Expand the main path and fold secondary capabilities.\n- Each node should keep only name, role, and key constraint or output.\n- Use code-like tokens for technical terms, such as `tool_call`, `final_output`, `SQLite`, or `prompts.jsonl`.\n- Avoid dense tiny tables.\n\nVisual hierarchy:\n- The controller, orchestrator, or core object has the strongest weight.\n- The main path is the clearest and most continuous line.\n- Output or terminal nodes must feel like a clear conclusion.\n- Decision or gate nodes must look like important judgment points.\n- Storage should feel grounded and quiet.\n- Observability or tracing should use low-contrast dashed lines.\n- Tools should feel callable, not equal to the main flow.\n\nConnection rules:\n- Sequential flow: strongest line\n- Branch / merge: secondary line\n- Loop: curved return line\n- Handoff: restrained highlight\n- Lookup / reference: thin line\n- Observation: low-contrast dashed line\n- Dependency: desaturated structural line\n- Bidirectional exchange: two-way connector\n- Traceability: fine dashed line\n\nHard bans:\n- Do not make a prettier Mermaid chart.\n- Do not preserve the original subgraph boxes by default.\n- Do not make every node the same size.\n- Do not add oversized icons.\n- Do not use rainbow colors, heavy shadows, glow, 3D, glassmorphism, or dense text.\n- Do not invent missing information.\n\nInput:\n{paste Mermaid source or diagram content here}\n</code></pre>\n","date_published":"2026-05-23T00:00:00.000Z","date_modified":"2026-07-31T00:00:00.000Z","tags":["AI","Prompting","Image Generation","gpt-image-2","Design"],"language":"en"},{"id":"https://www.lihuanyu.com/posts/gpt-image-2-commercial-cover-prompts/","url":"https://www.lihuanyu.com/posts/gpt-image-2-commercial-cover-prompts/","title":"gpt-image-2 提示词分享：商业封面与微缩视觉","summary":"整理适合 PPT 封面、报告插图、自媒体头图、科技行业封面和微缩视觉的 gpt-image-2 提示词模板。","content_html":"<p>商业视觉和内容配图之间有一条细线。太像广告，会油；太像资料图，又没有传播力。</p>\n<p>这一篇放的是偏封面、报告插图、行业视觉和品牌识别的 prompt。它们不只是要求模型“高级”，而是会安排主体、文字层级、信息密度和适用场景。</p>\n<p>这个系列里的模板主要来自 @xiaoxiaodong01 和 @MrLarus 的公开分享，在这里一并致谢。这里记录的是我用 <code>gpt-image-2</code> 跑过以后，觉得还值得复用的用法。</p>\n<p><a href=\"/en/posts/2026/gpt-image-2-commercial-cover-prompts/\">English version: gpt-image-2 Prompt Patterns: Commercial Covers And Miniature Visuals</a></p>\n<p>系列入口见 <a href=\"/posts/gpt-image-2-prompt-gallery/\">gpt-image-2 优秀提示词分享：可复用的生图模板</a>。</p>\n<h2>案例 1：信息视觉图像</h2>\n<p>有些内容不适合硬画成流程图。</p>\n<p>比如《岳阳楼记》。它不是一个数据报告，也没有什么现成的系统架构，但它有很强的空间感、情绪和层次。这个 prompt 有意思的地方，是把信息转成“微型景观”：平台、薄片、柱体、标签、留白和小人物，而不是硬凑一张图表。</p>\n<p>这张的留白和层次是对的，小字还是有一点 AI 字形感。正式使用前要检查文字，尤其是模型自己补出来的解释。图可以诗意，字不能乱来。</p>\n<p><img src=\"/assets/posts/2026/gpt-image2-prompts/02-info-visual.jpg\" alt=\"信息视觉图像示例：岳阳楼记\"></p>\n<p>生成备注：</p>\n<ul>\n<li>模型：<code>gpt-image-2</code></li>\n<li>尺寸：<code>1536x1024</code></li>\n<li>输出：JPEG，压缩后约 214 KB</li>\n<li>测试输入：岳阳楼记</li>\n<li>适合：PPT 封面、报告插图、知识内容封面、自媒体头图</li>\n</ul>\n<p>可复制模板：</p>\n<pre><code class=\"language-text\">请根据输入内容，创作一张具有高级编辑设计感的信息视觉图像。它不是普通图表，也不是模板化信息图，而是把信息、数据、观点、文字与情绪转化成一个明亮、克制、带有空间感的微型景象。\n\n在视觉创作之前，先判断信息来源。若输入中已经包含表格、截图、报告、图片、文字资料或明确数据，请优先使用这些材料，提炼其中最适合视觉化的结构。若只提供了主题、年份、行业、问题或方向，而没有足够数据，请先补充可靠资料并交叉核实；凡是涉及当前年份、未来年份、市场变化、平台趋势、价格、政策、榜单、人物、机构、品牌、新闻或任何可能随时间改变的信息，都不能凭空编造。若无法获得可靠资料，请先索要补充信息。画面可以诗意，但数据必须诚实。\n\n请把信息变成一组可以被观看的几何物件：立方体、细长柱体、薄片、切面、平台、容器、边界、悬浮构件、折叠结构、微小标记与空间层级。重要信息可以拥有更清晰的体量、更高的位置、更稳定的结构；次要信息可以更轻、更低、更贴近背景；隐性信息可以藏在边缘、空隙、折面、阴影、标签和细线之中。不要让画面像软件自动生成的图表，而要像一组被安静摆放、缓慢生长出来的信息景观。\n\n构图以大量留白为基础。主体可以居中，也可以轻微偏离中心，让空白成为画面的呼吸。画面需要一个稳定但轻盈的视觉承托，它不应是沉重的大黑块，而应像某种浅色结构、薄层平台、透明边界、漂浮底座或抽象器物。其他信息元素围绕它出现、穿过它、被它轻微遮挡，或从它的上下两侧延展。整体保留轻微透视与体积感，但不要变成真实 3D 渲染；保持平面插画、纸面印刷、低多边形体块和编辑设计之间的暧昧质感。\n\n配色不要通过列举具体颜色来决定，也不要使用固定色卡、流行模板或预设风格。请先根据主题的情绪、信息密度、使用场景和留白比例，生成一套属于这张图自己的色彩关系。色彩的核心不是选某几个颜色，而是建立明度、纯度、温度、重量和距离之间的秩序。整体应明亮、通透、干净，有空气感；背景像有光的纸面一样承托主体，主体色之间依靠明暗层次、冷暖偏移、透明叠加、面积大小和空间远近来区分信息层级。暗色只能作为极少量文字、细线、边缘、刻度、局部阴影或视觉停顿存在，不允许形成大面积压暗。不要让画面变脏、变闷、变厚重，也不要让色彩过度甜腻、商业化、荧光化或模板化。每一张图都应根据输入内容重新生成独立色系。\n\n文字不是说明书，而是画面结构的一部分。标题、短句、数字、百分比、标签、注释应自然嵌入几何体、边缘、空白和视线流动中。重要文字可以独立占据一片留白，像一个问题、一句判断或一句旁白；次要文字可以贴着柱体、折面、侧边、淡阴影或细小构件排列。中文可以竖排、横排、错位、贴边、悬停，但必须保持呼吸感。不要让文字填满画面，也不要让文字变成装饰噪音。它应该在需要被看见的时候出现，在不需要解释的时候退后。\n\n画面中可以出现少量微型人物。人物不需要复杂表情，可以只是站立、观察、搬运、测量、仰望、经过或停留。他们的存在是为了让信息拥有尺度，让宏大的数据变得有人味。人物也可以被简化成剪影、纸片、几何小像或极小的动作痕迹。除此之外，可以根据主题自然生成象征物，但它们必须被几何化、简化、安静化，不要变成素材堆砌。\n\n最终画面应适合用于 PPT 封面、报告插图、信息图、自媒体封面、品牌视觉、名片或知识内容展示。它应该既能承载信息，又像一件可以停留观看的视觉物。要有宏观的空白、微观的细节、克制的秩序、明亮的空气感、轻微的幽默，以及一种没有说满的余味。\n\n现在请用户输入：主题、核心表达、已有文字或数据、是否需要补充资料、是否有参考图片、希望使用的场景。\n用户输入的主题是：{请输入主题}\n</code></pre>\n<h2>案例 2：电子元器件微缩图文海报</h2>\n<p>这条我一开始觉得会翻车。</p>\n<p>“太阳就站在黑夜的门口”这种短句，太容易被模型画成廉价的光影海报。换成电子元器件微缩空间以后，反而多了一层约束：门、光、连接、等待，都要落到插座、开关、线缆、端子和微型人物的关系里。</p>\n<p>这个 prompt 适合科技行业封面、PPT 首页、展会视觉或者品牌主视觉。它真正有用的地方，是要求文字成为画面主角之一，不能被场景吞掉。</p>\n<p><img src=\"/assets/posts/2026/gpt-image2-prompts/03-electronics-poster.jpg\" alt=\"电子元器件微缩图文海报示例：太阳就站在黑夜的门口\"></p>\n<p>生成备注：</p>\n<ul>\n<li>模型：<code>gpt-image-2</code></li>\n<li>尺寸：<code>1536x1024</code></li>\n<li>输出：JPEG，压缩后约 226 KB</li>\n<li>测试输入：太阳就站在黑夜的门口</li>\n<li>适合：自媒体封面、PPT 首页、品牌主视觉、展会海报、社媒配图</li>\n</ul>\n<p>可复制模板：</p>\n<pre><code class=\"language-text\">请将用户输入转译为一张“电子元器件微缩图文海报”。\n\n先判断用户真正想完成的传播任务：是封面、信息图、PPT 首页、品牌主视觉、产品说明、展会海报、社媒配图，还是名片背景。画面必须既好看，也能被阅读。若用户输入了文字，文字就是画面的主角之一，不是小装饰。\n\n文字策略必须严格基于用户输入，不得自行添加与主题无关的口号、虚假数据、空泛行业词或莫名其妙的英文。若用户给了完整文案，只做排版和层级优化，不乱改含义；若用户只给关键词，可补充少量贴合主题的短标签或副标题；若用户没有给文字，则只生成高度概括、克制、准确的主题性文字。所有补充文字都要像从主题里自然长出来，而不是模板套话。\n\n请根据用途决定文字存在感：自媒体封面、海报、PPT 首页的核心关键词必须一眼可见，远看能捕捉；信息图需要清晰分区、标题、短标签、引导线和阅读顺序；品牌图或名片则应更克制、更有识别感。关键文字不能太小，不能藏在角落，不能被场景吞没。主标题、副标题、标签、注释之间要有明确层级。\n\n文字设计要和画面风格融合。它可以出现在建筑立面、展板、微型灯箱、站牌、地面导视、透明信息层、产品标签、边缘留白或 callout 标注中，也可以作为画面主标题形成强排版。字体应随主题变化：科技、工程、温暖、活泼、高级、工业、展会、社媒冲击感都可自然切换。中文要清晰，英文要简洁，字重、间距、对齐、留白要像专业设计。\n\n字体和信息组件可以带一点细小巧思：微型电流线、端子圆点、插孔形状、螺丝纹理、细线节点、标签折角、微光边框、刻度、编号、箭头、极小图标等。它们只能增强节奏和趣味，不能喧宾夺主，也不能变成廉价贴纸。\n\n画面由插座、排插、开关、断路器、线缆、端子、继电器、电源模块、插头、接线结构等真实电子元器件构成微缩空间。它们可以成为建筑、街区、车站、工坊、展台、塔楼、轨道、道路、桥梁或运行中的系统。元件要保持真实材质和可识别结构，不要变形成奇怪物体。\n\n构图必须清楚：一个主视觉核心，一个文字阅读区，一个人物叙事区，一个辅助信息区。可采用中心式、左右分栏、上下分层、斜向动线、留白标题、封面大字或信息图模块构图，但要根据用户意图自然选择。画面要有呼吸感，不能平均摆放，不能堆满。\n\n微型人物应成为叙事线索。他们可以检修、搭建、搬运、展示、协作、通勤、观察、排队或围绕核心装置工作。人物动作要小而准确。植物、轨道、铜线、道路、台阶、工具和标识牌可自然出现，用来引导视线、缓和工业感。\n\n请从用户主题中推导视觉隐喻：增长、连接、安全、能源、效率、协作、品牌、服务、制造、供应链等，都应通过空间关系、人物行为、元件结构和文字层级表达，不要直白解释。\n\n色彩不要套模板。根据主题、行业、情绪和用途智能生成配色。可以干净温暖、冷静科技、明亮社媒、高级克制、工业理性或轻快生活化。颜色必须服务阅读，保证文字与背景有足够对比。避免过度炫彩。\n\n整体呈现微缩模型摄影、产品广告摄影、浅景深、3/4 俯视角、柔和棚拍光、精密塑料与金属质感。最终作品应像一张真实拍摄的电子行业视觉海报，既有宏观秩序，也有局部细节。\n\n避免：真实品牌 Logo、乱码文字、无关文案、虚假数据、空泛口号、关键字过小、文字被遮挡、纯场景无信息层级、拥挤无留白、脏乱电路板、赛博霓虹、废土感、廉价玩具感、不可识别元件、比例过大的小人。\n\n请基于以下输入，自行判断画幅、用途优先级、文字主次、构图方式、配色气质、元件选择、人物行为、信息密度和视觉隐喻：\n\n用户输入内容：\n「{请输入一句话、主题或短文案}」\n</code></pre>\n<h2>案例 3：观念叙事型极简 Logo</h2>\n<p>Logo 类 prompt 很容易跑偏。</p>\n<p>模型最擅长给出“像 Logo 的图片”，但真正能用的标识，往往不是一个漂亮图案，而是一组可以被反复缩放、识别和延展的关系。这条 prompt 的价值在于，它没有直接要求“做一个高级 Logo”，而是要求图形有隐喻、有留白、有文字关系，整体像一个小型品牌识别装置。</p>\n<p>我用“墨屿 InkIsle”做测试。成图里一滴墨形成小岛，岛上立着一张像 Markdown 文档的白色页面，旁边有细线轨迹和一个很小的红点。<code>InkIsle</code> 拼写也没有出错。它还不是可以直接替换 favicon 的最终矢量稿，但作为品牌方向提案，已经能说明“墨、岛、写作、发布系统”这几个意象。</p>\n<p><img src=\"/assets/posts/2026/gpt-image2-prompts/12-inkisle-concept-logo.jpg\" alt=\"观念叙事型极简 Logo 示例：墨屿 InkIsle\"></p>\n<p>生成备注：</p>\n<ul>\n<li>模型：<code>gpt-image-2</code></li>\n<li>尺寸：<code>1024x1024</code></li>\n<li>输出：JPEG，压缩后约 16 KB</li>\n<li>测试输入：墨屿 InkIsle</li>\n<li>适合：品牌概念稿、创意工作室 Logo 方向、艺术项目识别、小众软件品牌提案</li>\n</ul>\n<p>可复制模板：</p>\n<pre><code class=\"language-text\">请根据用户输入的信息设计一个高完成度的「观念叙事型极简 Logo」。\n\n【用户输入】\n品牌名 / 项目名：【品牌名 / 项目名】\n副标题 / 产品名：【副标题 / 产品名】\n类型 / 行业：【设计工作室 / 影像工作室 / 创意厂牌 / 艺术展览 / 独立服饰 / 音乐品牌 / 建筑空间 / 创意出版 / 文化项目 / 小众品牌等】\n品牌定位：【品牌定位】\n核心概念：【梦境 / 控制 / 距离 / 孤独 / 探索 / 未来 / 想象 / 对话 / 结构 / 自由 / 叙事 / 精神性等】\n隐喻意象：【宇航员 / 兔耳朵 / 剪影 / 木偶 / 手 / 盒子 / 星球 / 门 / 眼睛 / 绳线 / 月亮 / 装置 / 几何结构 / 人物与空间关系等】\n情绪气质：【克制、冷静、先锋、神秘、实验、疏离、艺术感、诗性、理性、未来感等】\n主色调：【黑 / 白 / 灰 / 深褐 / 墨黑等】\n辅助色：【少量红色 / 暗红 / 极少量强调色】\n画幅比例：【1:1】\n\n【核心目标】\n要设计的不是普通企业 Logo，也不是单纯图形标，而是一个具有“观念表达、叙事张力和极简气质”的品牌标志。它需要通过极少的元素，传达一个概念、一种情绪、一段隐喻关系，形成一个像“小型视觉作品”一样的 Logo。\n\n最终效果应当具备：\n1. 极简但不空洞；\n2. 留白很多但不单薄；\n3. 图形很少但有意味；\n4. 文字克制但有设计感；\n5. 整体像独立设计师品牌、创意工作室或艺术项目的识别符号。\n\n【设计本质】\n这类 Logo 的核心不是直接“说明品牌做什么”，而是通过图形隐喻、场景关系和视觉诗意，去间接表达品牌的精神气质。\n请优先考虑：\n- 一个图形是否有观念；\n- 图形与文字之间是否形成一种叙事关系；\n- 留白是否增强了作品感；\n- 整体是否像一件小型艺术化识别装置。\n\n【最重要的原则】\n1. 图形必须简洁，但不能普通；\n2. 图形必须带有隐喻、象征或观念性；\n3. 不要把画面塞满，必须保持大量留白；\n4. Logo 主体可以偏小，像被放置在安静空间中的一个小型符号；\n5. 文字必须克制、简洁、有排版意识；\n6. 不要做成传统厚重商业徽章；\n7. 不要做成热闹卡通 Logo；\n8. 不要做成普通几何极简标；\n9. 整体要有实验感、先锋感、艺术感和独立品牌气质。\n\n【图形设计要求】\n请围绕【核心概念】和【隐喻意象】，设计一个“具有叙事意味的小型主图形 / 小场景 / 小装置”。\n\n图形要求：\n1. 不能只是普通插画；\n2. 不能只是常规图标；\n3. 需要有隐喻感；\n4. 需要有轻微叙事性；\n5. 可以有黑白剪影、细线结构、几何框架、少量红线、局部留白等形式；\n6. 图形整体应简洁，不宜太复杂；\n7. 图形要有“被观看”的价值，而不是仅仅功能化。\n\n【文字设计要求】\n文字部分应极其克制，但不能敷衍。\n要求：\n1. 品牌名 / 项目名清晰可读；\n2. 英文可偏简洁、细线、全大写、现代感；\n3. 中文可偏理性、冷静、少量使用；\n4. 文字可作为图形的注释、对位、分列、标题或支撑结构；\n5. 允许中英混排，但整体必须克制；\n6. 不要使用过于花哨的字体；\n7. 不要让文字成为主视觉堆积；\n8. 文字与图形之间要形成冷静的设计关系，而不是随便摆放。\n\n【排版与构图要求】\n整体构图应偏“展览感 / 提案感 / 作品集感”：\n1. 大量留白；\n2. 主体较小；\n3. 可以偏居中，也可以偏左 / 偏右，但要有设计平衡；\n4. 图与字的关系可分列、对置、上下悬置或横向对位；\n5. 可加入极细线、极少量小标记、小编号、小副标题；\n6. 画面不能像海报，也不能像包装正稿，而应像一个独立展示的品牌识别方案。\n\n【色彩要求】\n色彩要克制。\n推荐方案：\n- 黑白灰为主体；\n- 辅助以少量红色、暗红色、深褐色或一点点冷灰蓝；\n- 红色可以作为：切线、轨迹、连接线、警示点、局部结构线、强调关系；\n- 不要使用高饱和多彩配色；\n- 不要大面积渐变；\n- 不要热闹。\n\n【适用气质】\n整体应呈现：\n- 先锋\n- 冷静\n- 神秘\n- 诗性\n- 概念性\n- 设计师感\n- 小众品牌感\n- 独立工作室感\n\n【风格关键词】\n观念叙事型极简 Logo、Conceptual Narrative Minimal Logo、experimental logo、symbolic minimal logo、narrative mark、art-directed logo、independent studio branding、editorial-style brand mark、concept-driven brand identity、poetic visual metaphor。\n\n【验收标准】\n请确保最终结果满足：\n1. 图形很少但有意味；\n2. 画面很空但不单薄；\n3. 有一个清晰的观念或隐喻；\n4. 品牌名字体克制但有设计感；\n5. 整体像一件小型品牌视觉作品；\n6. 带有独立、先锋、设计师气质；\n7. 与传统商业 Logo 有明显区别；\n8. 适合创意、艺术、实验性品牌场景使用。\n\n【输出要求】\n请最终输出一个高完成度的「观念叙事型极简 Logo」，强调隐喻、叙事、小场景、极简构图、大留白、冷静文字与小众先锋气质。\n</code></pre>\n<h2>案例 4：东方留白感花卉知识图鉴</h2>\n<p>这条 prompt 适合做一种很轻的知识图鉴。</p>\n<p>它不是传统植物百科，也不是硬排版的信息图。它更像把一枝植物放在纸面上，再让少量知识点沿着枝干和留白自然停下来。信息有，但不急；画面有主体，但不压人。</p>\n<p>我用“蓝雪花”测试，要求只保留 6 个知识点、中心构图、超大主体和 2:1 横版。成图里蓝雪花的主体足够大，左侧标题、右侧和中部的知识点也都清楚。最适合放在 PPT、知识卡片、封面或花卉图鉴里。正式用时仍然要校对小字，比如学名、花期和养护要以可靠资料为准。</p>\n<p><img src=\"/assets/posts/2026/gpt-image2-prompts/13-blue-plumbago-knowledge-card.jpg\" alt=\"东方留白感花卉知识图鉴示例：蓝雪花\"></p>\n<p>生成备注：</p>\n<ul>\n<li>模型：<code>gpt-image-2</code></li>\n<li>尺寸：<code>2048x1024</code></li>\n<li>输出：JPEG，压缩后约 229 KB</li>\n<li>测试输入：蓝雪花</li>\n<li>适合：花卉知识图鉴、PPT 封面、知识卡片、自媒体封面、品牌视觉</li>\n</ul>\n<p>可复制模板：</p>\n<pre><code class=\"language-text\">请以一种克制、清透、东方留白感的视觉语言，创作一张适用于 PPT、信息图、自媒体封面、知识卡片、名片或品牌视觉的画面。画面不追求装饰堆叠，而像一件安静生成的作品：信息被轻轻托起，空间有呼吸，视觉元素彼此避让又彼此牵引。\n\n整体配色必须遵循低饱和、浅色、柔雾、近乎纸面的高级感。背景以极浅灰粉、暖白、淡米灰、雾白为主，不使用纯白死板背景；局部出现一块带有空气感的浅薄荷绿、冰蓝绿、淡青色或水雾色块，它不是标准硬边几何，而像被光稀释后的半透明平面。文字颜色使用灰黑、烟灰、深褐灰，不使用高饱和强色。点缀色可以来自自然物：嫩芽绿、枝干褐、花瓣白、浅粉，但必须轻、少、准。\n\n构图采用非对称留白。主体不要居中压满，而是在画面中形成斜向生长、纵向呼吸或轻微偏移的关系。可以有一条自然枝干、柔软线条、植物、纤细曲线、云雾状形体、纸面纹理、淡淡投影或抽象数据轨迹作为视觉主线。它们像自然生长出来，而不是机械摆放。图形要轻盈、柔软、带微弱透明感和虚实层次，避免硬朗科技线框、标准图标模板、商业素材拼贴感。\n\n文字排版是画面的骨架。标题可以纵排、竖向分布、分段留白，字距要疏朗，字号有节制但有仪式感。英文或辅助信息使用极细无衬线、大字距、小字号，像空气中的标注。正文信息不要堆满，适合被拆成短句、诗性短行、数据标签或信息节点。重要数字可以放大，但仍保持安静、优雅，不做电商式冲击。文字与图形之间要有互相成全的关系：文字像落在空间里的秩序，图像像托住信息的气息。\n\n若用户输入的是知识、数据、观点或报告内容，请将其转化为一种轻量信息图：用少量线段、浅色块、微型标签、数字层级、纵向节奏、留白分组来组织信息。不要做传统表格，不要做密集流程图，不要做机械仪表盘。信息应该像被自然地安置在画面中，每一组文字都有自己的停顿和位置。\n\n画面光感柔和，低对比，自然漫射光，带轻微景深与柔焦。可以有极淡阴影、纸张肌理、空气颗粒、半透明叠层。所有元素边缘都要克制，不要锐利过度，不要霓虹，不要赛博，不要高饱和渐变，不要厚重 3D，不要卡通，不要插画感过强。\n\n最终画面应像同一位作者延展出的系列作品：安静、清浅、秩序感强，却不死板；信息明确，却不喧哗；有东方节气般的自然意象，也有现代信息设计的理性骨架。\n\n请根据以下用户变量生成画面：\n\n主题：\n{请输入主题}\n\n用途：{请输入用途}\n\n文字知识点：{请输入知识点数量和内容要求}\n\n构图要求：{请输入构图、主体大小和画面关系}\n\n比例：{请输入比例}\n</code></pre>\n<h2>什么时候用</h2>\n<p>如果一张图要放在 PPT 首页、报告封面、自媒体头图、活动视觉、品牌方向提案或知识卡片里，可以先从这里挑。正式发布前重点检查两件事：主标题是否可读，模型补的小字有没有乱编。Logo 类图片还要单独做矢量化和小尺寸测试，不能只看大图好不好看。</p>\n","date_published":"2026-05-23T00:00:00.000Z","tags":["AI","提示词","生图","gpt-image-2","设计"],"language":"zh"},{"id":"https://www.lihuanyu.com/posts/gpt-image-2-city-wallpaper-prompts/","url":"https://www.lihuanyu.com/posts/gpt-image-2-city-wallpaper-prompts/","title":"gpt-image-2 提示词分享：城市海报与艺术壁纸","summary":"整理适合城市收藏海报、旅行视觉、拼音主题海报和艺术壁纸的 gpt-image-2 提示词模板。","content_html":"<p>有些图不需要解释太多，耐看更重要。</p>\n<p>这一篇放的是偏壁纸、城市海报和收藏视觉的 prompt。它们的重点不是塞信息，而是控制留白、色彩、构图和长期观看的舒服程度。</p>\n<p>这个系列里的模板主要来自 @xiaoxiaodong01 和 @MrLarus 的公开分享，在这里一并致谢。这里记录的是我用 <code>gpt-image-2</code> 跑过以后，觉得还值得复用的用法。</p>\n<p>系列入口见 <a href=\"/posts/gpt-image-2-prompt-gallery/\">gpt-image-2 优秀提示词分享：可复用的生图模板</a>。</p>\n<h2>案例 1：清冷油墨风艺术壁纸</h2>\n<p>这条和前面几条不太一样，它几乎不承担信息说明任务。</p>\n<p>我用“星际穿越中的经典壁纸级画面”做测试。它不需要模型解释剧情，也不需要生成大段文字，只要把黑洞、海面、小人物和孤独感放到一张耐看的图里。</p>\n<p>这类 prompt 的价值在于克制。它反复强调留白、旧感、雾感、手工痕迹和“不要解释太满”。成图的黑洞和海面都有辨识度，小人物也没有抢戏，比较适合当横版壁纸。</p>\n<p><img src=\"/assets/posts/2026/gpt-image2-prompts/05-cold-ink-wallpaper.jpg\" alt=\"清冷油墨风艺术壁纸示例：星际穿越中的经典壁纸级画面\"></p>\n<p>生成备注：</p>\n<ul>\n<li>模型：<code>gpt-image-2</code></li>\n<li>尺寸：<code>1536x1024</code></li>\n<li>输出：JPEG，压缩后约 366 KB</li>\n<li>测试输入：星际穿越中的经典壁纸级画面</li>\n<li>适合：桌面壁纸、手机壁纸、艺术插图、情绪封面</li>\n</ul>\n<p>可复制模板：</p>\n<pre><code class=\"language-text\">请根据用户最后输入的【主题】或【参考图片】，生成一张具有强艺术家气质、强审美判断、强风格表达的单幅艺术绘画图像。\n\n你需要先判断输入类型，并执行对应逻辑：\n\n- 如果用户输入的是【主题 / 关键词 / 短句 / 概念】：先理解其气质、情绪、象征、联想与隐藏张力，再从“艺术家的视角”进行视觉转译，而不是直接图解主题。\n- 如果用户提供的是【一张或多张参考图片】：不要直接复制原图，而是提炼其中最有价值的主体、姿态、构图重心、物象关系与情绪气息，再以更有艺术判断的方式重新创作。\n- 如果用户同时提供【主题 + 图片】：以图片内容为视觉基础，以主题作为情绪和表达方向，让最终画面既保留原图核心识别度，又完成真正有灵气的风格化重构。\n\n这不是普通插画，不是通用唯美风格图，也不是简单滤镜转换。它必须像一位有审美、有手感、有临场判断的顶级艺术家创作出来的作品。画面应有灵气、有呼吸感、有微妙的失衡感和恰到好处的“非标准处理”，而不是机械、均匀、过分工整。\n\n【核心气质】\n整体应轻、柔、静、透、旧、雾、松，带有明显的孤独感、安静感与凝视感。像一张被空气、纸纹、时间和柔光轻轻覆盖过的艺术绘画。带有粉彩、油画棒、干湿混合颜料、粗纹画布、磨砂玻璃般的朦胧质感。不要高清，不要锐利，不要塑料感，不要数码感。\n\n【大量留白 / 孤独气质】\n画面必须具备高质量留白，不要塞满，不要过度充实，不要把主体撑满画面。留白不是空，而是情绪的一部分，是孤独、距离、呼吸感和高级感的重要来源。\n应主动使用大面积安静背景、空旷区域、柔和过渡区域，让主体在留白中被凝视、被衬托、被放大情绪。\n整体气质应有轻微的孤独感、疏离感、静默感，像一个被安静观看的瞬间，像情绪停留在空间里，而不是热闹的叙事画面。\n\n【顶级艺术家构图逻辑】\n构图必须是一张完整的单幅画面，不要分区，不要四宫格，不要拼贴。构图水平必须体现出顶级艺术家的审美判断，不要普通，不要平铺直叙，不要只是把主体规规矩矩摆在正中间。\n要具备真正的高级构图意识，可灵活使用：\n- 大量留白\n- 非对称平衡\n- 偏置重心\n- 边缘裁切\n- 局部放大\n- 远近反差\n- 视觉停顿\n- 暧昧焦点\n- 主体偏角落或偏一侧的克制处理\n- 局部遮挡\n- 静物化凝固\n- 带有策展意识的画面秩序\n\n重点不是“完整交代主题”，而是“让画面本身有艺术张力、节奏、余味和记忆点”。\n\n【壁纸适配要求】\n最终图像必须达到适合做用户桌面壁纸和手机壁纸的水平。构图要耐看、耐久看，不廉价，不花哨，不堆信息。\n应注意以下几点：\n- 画面整体要干净、耐看、统一，具有长期观看价值\n- 不要让主体把画面中心全部占满，需保留适当空白区域\n- 预留适合壁纸使用的呼吸区与空域，让画面在桌面图标、手机时间区域存在时依然好看\n- 视觉重心可适度偏左、偏右、偏下或偏角落，形成更高级的壁纸构图\n- 画面需要兼顾横版与竖版审美逻辑：既像一张高端艺术作品，也像一张真正想保存下来的壁纸\n- 不要做成信息图，不要做成大段排版，不要做成海报化说明图，而是一张纯粹的艺术图像\n\n【真正的配色逻辑】\n不要把颜色理解成“统一低饱和”或“均匀柔和”。真正的色彩语言应是：\n\n1. 以安静、偏冷、偏灰的底色统摄全画面：\n灰蓝、蓝绿、湖水绿、雾青、浅青灰、冷灰白、淡奶灰等作为基础氛围色，用来建立安静、湿润、柔雾般的空间感。\n\n在局部突然出现少量非常巧妙的彩色亮点：\n这些亮点色不是平均分布，也不是规则点缀，而像艺术家在某一处凭感觉“突然加进去的一笔妙色”。\n可灵活出现：珊瑚橙、蜜桃粉、番茄红、柠檬黄、奶油黄、明亮草绿、钴蓝、浅紫、亮白、暖橘等。\n\n这些彩色必须“出现得巧妙”：\n它们通常只出现在局部关键位置，例如眼睛、边缘、转折处、反光处、花蕊、果实、叶尖、器物的一小块表面、主体某个最值得凝视的细节。\n不是为了装饰，而是为了让画面突然活一下、亮一下、灵一下。\n\n允许颜色出现“不完全合理但很好看”的关系：\n比如一个物体，不必完全遵守真实固有色，可以在灰绿中揉入橙黄、蓝紫、奶白、淡粉、亮绿，形成一种艺术化混色与局部跳色。\n这种颜色关系可以带一点意外感、手工感、灵感感，而不是教科书式上色。\n\n色彩不是平涂，而要带有“擦、抹、糊、蹭、渗、压、覆盖”的痕迹：\n颜色之间可以互相侵入、染开、叠压、局部涂抹，形成一种生动的、不那么规矩的绘画性。\n局部可以出现像被手指、干刷、粉笔、布面擦开的痕迹，让某些颜色不是“填进去”，而是“抹出来”的。\n\n【艺术表达手法】\n画面必须具有艺术家的处理感，而不是机械生成感。请主动使用以下视觉方法：\n- 某些局部故意不画完整，而是半完成、半溶解、半擦除\n- 主体表面可出现局部涂抹、蹭色、覆盖、拖拽、断裂边缘\n- 某一块颜色可以像被随手抹开，形成非常自然的绘画性痕迹\n- 某些轮廓不必封闭，不必太干净，可轻微松散、破碎、化开\n- 允许局部存在“模糊但很对”的笔触，而不是处处清楚\n- 可保留一些像画家临时改变主意的痕迹，让画面更有生命力\n- 画面不需要把一切解释清楚，而应保留暧昧、余味、呼吸和想象空间\n\n【主题 / 图片转译逻辑】\n无论输入的是主题还是图片，都不要只做表面复现。应自动寻找最值得被凝视的视觉瞬间，把主体进行审美提炼。\n主体可以是局部、特写、侧面、轮廓、切片、漂浮物、静物、象征物或抽象意象。重要的不是“说明”，而是“表达”。\n\n【如果用户提供参考图片】\n请保留以下内容：\n- 主体的基本识别度\n- 核心构图关系\n- 主要情绪氛围\n- 最重要的视觉重心\n\n但同时主动进行艺术化处理：\n- 删掉杂乱细节\n- 简化背景\n- 提炼形体\n- 加入更有判断的配色变化\n- 增加局部擦抹、覆盖、化开、断裂、混色等绘画性处理\n- 不要做普通照片转插画，而要做真正的艺术重构\n\n【笔触与材质】\n使用粉彩、油画棒、干刷、厚薄不均的颜料痕迹、柔软的色粉感、略带粗糙的布面纹理。保留颗粒、擦痕、刷痕、叠色、模糊高光、低对比阴影和轻微失焦边缘。\n边缘不要太利落，颜色可以轻轻渗进去，像停留在粗布和纸面里。\n\n【光线】\n柔和漫射光，像清晨窗边、阴天室内、旧画室里的自然光。不要摄影棚光，不要赛博光，不要硬阴影，不要3D渲染感。\n\n【最终目标】\n最终生成的不是普通“主题插画”或“照片转绘”，而是一张真正具有艺术家视角、配色灵气、手工痕迹、审美判断、孤独气质、大量留白、壁纸级构图与高级记忆点的单幅艺术绘画作品。\n它应当安静但不乏味，柔和但不灰闷，带彩但不艳俗，松动但不潦草，空旷但不空洞，艺术化但不做作。\n\n【禁止】\n不要水印、logo、网址、署名\n不要真实摄影感\n不要平滑数字插画感\n不要机械平均配色\n不要所有颜色都很闷\n不要普通居中摆拍\n不要廉价唯美风\n不要过度完整、过度解释\n不要明显 AI 塑料感\n不要画面拥挤\n不要信息太满\n不要失去留白与孤独感\n\n用户输入内容：{请输入主题、关键词、短句或参考图片}\n比例：{4:3 / 16:9 / 9:16 / 1:1}\n</code></pre>\n<h2>案例 2：极简线条城市海报</h2>\n<p>城市海报很容易画成明信片。</p>\n<p>明信片不是不好，只是太直白：一个地标，一个大标题，再加一点天空和水面。这个 prompt 更强调“真实城市切面”，要把建筑、街道、标牌、人物活动和当地文字系统放在一起。</p>\n<p>我用“上海外滩”测试。成图里的竖版构图、建筑线描和路牌文字都比较稳，外滩历史建筑和陆家嘴天际线也放进了同一张画面里。严格说它仍然有一点理想化，但已经比普通旅游宣传图更像一张可以收藏的城市海报。</p>\n<p><img src=\"/assets/posts/2026/gpt-image2-prompts/08-shanghai-line-poster.jpg\" alt=\"极简线条城市海报示例：上海外滩\"></p>\n<p>生成备注：</p>\n<ul>\n<li>模型：<code>gpt-image-2</code></li>\n<li>尺寸：<code>1024x1536</code></li>\n<li>输出：JPEG，压缩后约 257 KB</li>\n<li>测试输入：上海外滩</li>\n<li>适合：城市海报、旅行收藏图、品牌视觉、空间装饰画</li>\n</ul>\n<p>可复制模板：</p>\n<pre><code class=\"language-text\">请根据最后输入的主题创作一幅超高分辨率竖版极简线条城市海报。\n\n如果主题是一座城市，请先理解这座城市的核心气质，从空间景观、街头生活、人文风情、商业氛围、饮食记忆、地方文化、历史层次与日常节奏中提炼视觉线索，并将这些线索自然组织进同一幅完整而统一的城市画面之中，使画面呈现真实而鲜活的城市切面，而不是零散拼贴。\n\n如果主题是具体街道、街区、地标、商圈或代表性地点，请以该地点的真实空间关系为核心，准确呈现街道结构、建筑界面、店铺招牌、交通设施、文字系统、人群状态与现场氛围，并保留其最有辨识度的在地特征。\n\n如果主题包含地址、坐标或参考图片，请优先依据真实信息理解场景，不要凭空杜撰，不要随意替换地点特征，不要画成与实际地点无关的通用城市景观。\n\n整体画面应避免廉价旅游宣传式表达，不做明信片式直白展示，而是捕捉一个真实、精致、富有生活感的城市瞬间。让建筑、街道、标牌、橱窗、街头设施与人物活动自然共生，地标克制地融入环境，不夸张，不喧宾夺主。人物应体现本地真实穿着、年龄层次、生活方式与行动节奏，可出现通勤、步行、交谈、骑行、购物、等候、用餐或休憩等自然状态。\n\n画面采用竖版正面街头视角，使用极简矢量线描、纤细准确的单线、清晰几何透视、克制留白与高密度但有秩序的细节组织，形成安静、现代、灵动且高级的视觉秩序。整体应具有城市品牌视觉与收藏级旅行海报的完成度。\n\n顶部设置醒目的主标题，副标题使用当地语言与国家或地区信息。所有街头文字、招牌与排版必须清晰、自然、真实、专业，避免乱码、错误文字与随意拼写。\n\n关于色彩，请不要从固定选项中机械挑选，也不要让不同城市反复落入相似的配色模板。请先理解主题本身的气候、光线、时间感、建筑材料、历史气息、产业结构、商业温度、饮食印象、自然环境与情绪质地，再由这些因素综合提炼出专属于该主题的主色与底色关系。每个主题的颜色都应是独立生成的，具有明确的在地性与审美依据，而不是套版结果。允许颜色呈现温度、湿度、年代感、都市能量或文化性格上的差异。主色负责建立城市的精神气质，底色负责提供空气感与纸感，两者共同构成统一而鲜明的单色印刷美学。即使只使用一组主色与底色，也应通过线条疏密、明度层次、局部压重、留白比例与细节节奏形成丰富而细腻的视觉变化，避免单调和平铺。\n\n最终成品需呈现超高分辨率、可打印、线条清晰、细节丰富、风格统一、审美高级的城市海报效果。\n\n主题：{请输入城市、街区、地标或具体地点}\n</code></pre>\n<h2>案例 3：城市拼音主题海报</h2>\n<p>这条最需要检查的是拼写。</p>\n<p>它会把中文主题转成英文或拼音罗马化标题，再让巨型字母变成主题橱窗。这个思路很好看，但只要标题错一个字母，整张图基本就废了。</p>\n<p>我用“杭州”测试，主标题是 <code>HANGZHOU</code>，没有明显错字。西湖、桥、塔、茶田和城市天际线也都被放进字母里。正式使用时，第一眼先看拼写，再看图好不好看。</p>\n<p><img src=\"/assets/posts/2026/gpt-image2-prompts/09-hangzhou-romanized-poster.jpg\" alt=\"城市拼音主题海报示例：杭州\"></p>\n<p>生成备注：</p>\n<ul>\n<li>模型：<code>gpt-image-2</code></li>\n<li>尺寸：<code>1536x1024</code></li>\n<li>输出：JPEG，压缩后约 232 KB</li>\n<li>测试输入：杭州</li>\n<li>适合：城市拼音海报、主题收藏海报、品牌视觉、长图封面</li>\n</ul>\n<p>可复制模板：</p>\n<pre><code class=\"language-text\">生成一张超高分辨率、印刷级别、具有收藏价值的主题海报。用户会在结尾输入一个主题词，它可以是城市、地点、人物、品牌、节日、文化概念、物件、情绪、建筑、自然元素或任意名词。请围绕这个主题词建立完整视觉系统，并将它转化为画面中央的巨型英文主标题；若主题词不是英文，请自然转换为准确、优雅、适合作为海报标题的英文表达或拼音罗马化形式，不要出现乱码、错误拼写或多余解释。\n\n画面以巨型主标题为核心，字母高耸、粗壮、清晰，像一组被打开的主题橱窗。每个字母之中延展出与主题相关的场景、符号、结构、纹理、空间与叙事片段，它们彼此连通，形成连续而完整的视觉故事。不要机械切割画面，不要平均分配内容，让每个字母之中都像一个被精心组织却仍然自然流动的世界。\n\n顶部设置一条横向剪影式装饰带，可根据主题自然生成相应元素：轮廓、图形、线条、符号、交通、建筑、植物、器物、抽象结构、程序化几何或其他相关视觉片段。它不是孤立装饰，而应与主体在色彩、节奏、密度和气质上完全统一，共同构成完整的版式张力。\n\n整体风格兼具高端编辑设计、中世纪现代、瑞士平面设计与扁平几何插画的气质。构图克制，留白讲究，线条准确，边缘利落，画面安静但具有吸引力。色彩逻辑必须建立在“明亮、柔和、低饱和”的基础上，整体气质清透、轻盈、优雅、精致，避免沉闷、厚重、灰暗或脏浊。以奶白、雾粉、浅蓝、薄荷绿、砂岩、浅灰这一类色相为基础组织画面，也可以延展出柔和杏色、鼠尾草绿、粉蜡橙、淡芥末、灰玫瑰、浅雾紫等相近体系，但所有颜色都应保持低饱和、高级、通透、有空气感。\n\n配色不要走单色压暗路线，也不要落入自然主义的蓝天白云默认方案。应采用一种更聪明的综合色彩编排：以浅亮底色为画面基底，以温柔而克制的综合色组建层次，再以少量更明确的强调色作为视觉锚点。强调色可以略微跳出，但仍需保持审美约束，像一枚点醒画面的按钮，而不是喧宾夺主。整体色彩关系应具有电影感与设计感，明亮但不刺眼，柔和但不寡淡，低饱和但不无聊，丰富但不杂乱，轻盈却不发飘，复古却不陈旧。\n\n暗部必须被严格控制，只作为局部结构、轮廓、节奏或视觉支点使用，不能大面积压低画面情绪。阴影应偏轻、偏薄、偏干净，不制造沉重戏剧感。整幅作品应更接近一种经过精密校准的明亮世界：温柔、清醒、聪明、讲究，并具有让人想收藏的视觉愉悦。\n\n所有可见文字必须使用英文或准确罗马化标题，拼写绝对正确，排版专业、稳定、干净，不出现乱码、变形、破碎字母、随机符号或多余文字。整体效果应达到博物馆级、收藏级海报水准，既有主题识别度，也有大师级的色彩判断与形式完成度。若无额外要求，默认画幅比例为 2:1。\n\n主题词：{请输入主题词}\n</code></pre>\n<h2>什么时候用</h2>\n<p>如果要做桌面壁纸、手机壁纸、城市旅行海报、空间装饰画或主题收藏图，可以看这一篇。城市类图先检查地名、拼写和地标关系；壁纸类图则看它能不能经得住多看几天。</p>\n","date_published":"2026-05-23T00:00:00.000Z","tags":["AI","提示词","生图","gpt-image-2","设计"],"language":"zh"},{"id":"https://www.lihuanyu.com/posts/gpt-image-2-technical-diagram-prompts/","url":"https://www.lihuanyu.com/posts/gpt-image-2-technical-diagram-prompts/","title":"gpt-image-2 提示词分享：技术图解与信息图","summary":"整理适合技术方案、架构说明、报告插图和 Mermaid 重构的 gpt-image-2 提示词模板。","content_html":"<p>技术内容最怕两件事：一是图里什么都有，读者什么都抓不住；二是图看起来很漂亮，但技术链路已经变形。</p>\n<p>这一篇只放偏技术说明的 prompt。它们的共同点是先压缩信息，再安排主路径、节点角色、线条层级和结论位置。图可以好看，但不能为了好看把机制讲错。</p>\n<p>这个系列里的模板主要来自 @xiaoxiaodong01 和 @MrLarus 的公开分享，在这里一并致谢。这里记录的是我用 <code>gpt-image-2</code> 跑过以后，觉得还值得复用的用法。</p>\n<p><a href=\"/en/posts/2026/gpt-image-2-technical-diagram-prompts/\">English version: gpt-image-2 Prompt Patterns: Technical Diagrams And Infographics</a></p>\n<p>系列入口见 <a href=\"/posts/gpt-image-2-prompt-gallery/\">gpt-image-2 优秀提示词分享：可复用的生图模板</a>。</p>\n<h2>案例 1：手绘知识图解</h2>\n<p>开发者最常缺的不是图，而是能把一堆机制讲清楚的图。</p>\n<p>我拿 “Hermes Agent 架构总览” 做测试，是因为这类内容很容易画成两种极端：要么像一张冷冰冰的流程图，要么把所有文字塞满画布。这个 prompt 的好处是，它会先要求模型压缩信息，再用模块、箭头、标签和底部结论建立阅读路径。</p>\n<p>成图里标题、主链路、<code>Guardrails / Observability</code> 和底部总结都还算清楚。用在正式技术文章里时，输入内容最好控制在 6 个模块以内。模块再多，模型也能画，只是中文小字开始变得不稳。</p>\n<p><img src=\"/assets/posts/2026/gpt-image2-prompts/01-handdrawn-knowledge.jpg\" alt=\"手绘知识图解示例：Hermes Agent 架构总览\"></p>\n<p>生成备注：</p>\n<ul>\n<li>模型：<code>gpt-image-2</code></li>\n<li>尺寸：<code>1536x1024</code></li>\n<li>输出：JPEG，压缩后约 359 KB</li>\n<li>测试输入：Hermes Agent 架构总览</li>\n<li>适合：技术方案图、架构总览、复盘汇报、知识卡片、团队分享配图</li>\n</ul>\n<p>测试输入大意如下：</p>\n<pre><code class=\"language-text\">Hermes Agent 架构总览\n\n核心判断：Hermes Agent 的关键不是把大模型包成聊天入口，而是把「任务理解、计划编排、工具调用、状态记忆、结果交付」拆成可观察、可替换、可复用的工程链路。\n\n主链路：\n用户请求 → Orchestrator 编排器 → Planner 任务拆解 → Tool Router 工具路由 → Tools / APIs 执行 → Memory / State 状态沉淀 → Response Builder 结果组织 → 用户确认或继续迭代\n\n关键模块：\n1. Orchestrator：接收请求，维持会话上下文，决定下一步交给哪个模块。\n2. Planner：把目标拆成步骤，识别依赖、风险和需要补充的信息。\n3. Tool Router：根据计划选择代码、搜索、文档、数据库、浏览器等工具。\n4. Memory / State：保存用户偏好、任务状态、中间产物和可复用上下文。\n5. Guardrails / Observability：记录 tool_call、错误、耗时、权限边界和回滚点。\n6. Response Builder：把执行结果整理成可交付答案、代码变更或下一步建议。\n\n底部结论：一个可维护的 Agent 架构，应当让每一次决策、每一次工具调用和每一次状态变化都能被追踪。\n</code></pre>\n<p>可复制模板：</p>\n<pre><code class=\"language-text\">请把我提供的内容转化成一张高可读性的手绘知识图解。风格像认真整理过的创意手帐 + 白板推演 + 咨询报告信息图，而不是冰冷模板。\n\n【输出目标】\n生成一张适合传播、汇报和复用的知识图解。它必须先让人抓住核心判断，再沿着模块逐步阅读，最后记住一句结论。\n\n【语言要求】\n图上所有可见文字根据用户的输入来确定语言，中文，英文或其他。\n不要混用语言，除非是技术名词、产品名、协议名、代码路径或数字指标。\n\n【画布要求】\n比例：{16:9 / 5:4 / 4:3 / 21:9}\n质量：4K high resolution\n背景：浅米白 / 浅暖灰，保留轻微纸张纹理和呼吸感。\n整体清晰、留白稳定，不要把文字挤到看不清。\n\n【信息设计规则】\n不要逐字搬运原文。先压缩信息，再画图。\n请把内容整理成：\n1. 顶部：强标题 + 一句话核心判断\n2. 中部：3–6 个主模块，按流程、对比、阶段或因果关系排列\n3. 模块内：每个模块最多 3–5 条短 bullet\n4. 底部：一条 Flow Summary / Decision Summary / Bottom Line\n5. 如果内容很多，只保留最关键的 8–10 个判断，避免微型文字\n\n【可读性规则】\n标题必须最大、清楚、有重量。\n模块标题要有秩序，正文必须短句化。\n每个模块不要超过 6 行正文。\n每条 bullet 尽量简短。\n不要使用密密麻麻的小字表格。\n不要为了完整而牺牲可读性。\n\n【视觉风格】\n黑色或深墨色手写线条建立阅读骨架。\n使用圆角分区、细线框、轻阴影、编号、箭头、标签和小图标。\n线条允许轻微手绘抖动，但整体对齐、边距、分组要稳定。\n图标只做路标和强调，不要抢走文字层级。\n\n【配色规则】\n使用克制的标记笔色彩：\n浅米白背景 + 黑色主线条；\n低饱和青绿、鼠尾草绿、淡紫、柔橙、浅蓝作为分区和路径颜色。\n避免霓虹色、强渐变、过度商业光效和整页单色化。\n彩色区域只占少量到中等面积。\n\n【准确性规则】\n严格保持输入内容中的技术链路、组件名称、箭头方向、协议、端口、数据流和判断。\n不要自行新增未提供的组件。\n不要把动作写错，例如“读取日志”不能画成“生成日志”。\n如果空间不足，优先保留主链路、关键差异和最终判断，删掉次要解释。\n\n【内容】\n{请输入你的内容或者参考图片}\n</code></pre>\n<h2>案例 2：Mermaid 信息图重构</h2>\n<p>Mermaid 很适合写在文档里，但不一定适合放在 PPT 或文章头图里。</p>\n<p>我测试的是一段 <code>imgasset</code> 图片流水线：写 <code>prompts.jsonl</code>，生成原图，压缩，上传，再放进 Markdown。原始 Mermaid 很清楚，但视觉上还是偏工程草图。</p>\n<p>这条 prompt 不只是“美化 Mermaid”。它会要求模型先理解语义，再重新组织主路径、节点角色、线条层级和视觉权重。复杂图最好先把输入压短一点。节点太多时，模型为了完整会塞小字，最后反而不如原图。</p>\n<p><img src=\"/assets/posts/2026/gpt-image2-prompts/07-mermaid-infographic.jpg\" alt=\"Mermaid 信息图重构示例：imgasset 图片流水线\"></p>\n<p>生成备注：</p>\n<ul>\n<li>模型：<code>gpt-image-2</code></li>\n<li>尺寸：<code>1536x1024</code></li>\n<li>输出：JPEG，压缩后约 125 KB</li>\n<li>测试输入：一段 <code>imgasset</code> 图片流水线 Mermaid</li>\n<li>适合：技术文档配图、架构讲解、流程说明、PPT 技术页</li>\n</ul>\n<p>测试输入：</p>\n<pre><code class=\"language-mermaid\">flowchart LR\n  A[写 prompts.jsonl] --&gt; B[imgasset generate]\n  B --&gt; C[raw 原图]\n  C --&gt; D[TinyPNG 压缩]\n  D --&gt; E[publish 图片]\n  E --&gt; F[CDN 上传]\n  F --&gt; G[Markdown 引用]\n</code></pre>\n<p>可复制模板：</p>\n<pre><code class=\"language-text\">你是高级信息图重构生成器，兼具信息架构师、视觉设计总监、技术机制图设计师能力。\n\n任务：\n将用户提供的 Mermaid / C4 / Flowchart / Sequence / State / ER / Timeline 源码，或其渲染图片，重新设计为一张高保真、高审美、专业级信息图。\n\n目标不是美化原图，不是复刻 Mermaid，不是普通流程图换皮。\n你必须先理解其语义结构，再重新编译为新的信息架构图。\n\nOutput:\n生成一张最终高保真信息图。\n不要输出分析过程、解释、Markdown、源码或设计说明。\n\nInput truth rules:\n1. 如果输入是源码，源码是语义真相。忽略 Mermaid 原始布局、颜色、classDef、节点样式。\n2. 如果输入是图片，图片只作为语义提取材料。不得临摹原图布局、配色、节点形状或箭头路径。\n3. 图片文字模糊时，只保留可确认信息，不要编造业务逻辑。\n\nSemantic extraction:\n提取：\n- entities\n- groups\n- actors\n- relationships\n- branches\n- merges\n- loops\n- gates\n- tools\n- stores\n- schemas\n- states\n- outputs\n- triggers\n- annotations\n- dependencies\n- observations\n\nRole assignment:\n为每个实体分配角色：\n- input\n- output\n- controller\n- orchestrator\n- processor\n- resolver\n- decision\n- gate\n- tool\n- storage\n- observer\n- actor\n- artifact\n- annotation\n- boundary\n- event\n- state\n- terminal\n- reference\n\nPrimary mechanism:\n必须识别唯一主机制，并让它在 3 秒内可见。\n主机制类型可为：\n- pipeline\n- orchestration\n- resolver pipeline\n- gating\n- handoff\n- layered system\n- lifecycle\n- hub-and-spoke\n- dependency network\n- sequence interaction\n- state transition\n- artifact-centered flow\n- split decision tree\n- release / deployment flow\n\nPrimary path:\n优先提取：\ninput → controller / processor / decision / resolver → output\n主路径必须成为主视觉轴。\n辅助关系只能作为分支、回路、引用线、观察线、traceability 线、dependency 线。\n\nTemplate selection:\n- Linear primary path with &lt;= 7 major steps → Core Flow Spine\n- Central controller with &gt;= 3 outgoing branches → Orchestrator Hub\n- Runtime / gateway / engine / tools / storage → Layered Blueprint\n- Validation / conditional branching → Split Gate\n- Multi-actor interaction → Swim Relay\n- User action + internal state toggle → Interaction State Panel\n- Artifact / version / tag resolves execution → Artifact Anchor Resolver\n- Entity relationship → Relational Data Grid\n- Dense dependency graph → Clustered Zones\n- Lifecycle / state transition → Lifecycle Ring\n\nStyle selection:\n- SDK / Agent / orchestration / migration / platform / architecture → Premium Technical Editorial\n- Business process / product flow / user decision → Premium Light\n- Documentation / explanatory mapping → Neutral Editorial\n- Layered infra / system runtime → Technical Blueprint\n- AI / security / real-time observability only → Dark Futuristic\nDefault style: Premium Technical Editorial\n\nCanvas defaults:\n- aspect ratio: 16:10\n- canvas target: 1600×1000\n- outer padding: 72\n- section gap: 56\n- node gap: 28–40\n- max visible nodes: 18\n- max primary path nodes: 7\n- max annotation groups: 3\n\nInformation compression:\n- 同类节点超过 4 个时合并为模块组\n- 主路径展开，辅助能力折叠\n- 每个节点只保留：名称 + 角色 + 关键约束 / 输出\n- 长文本压缩为短标签\n- 技术词使用 code token，例如 `tool_call`, `final_output`, `SQLite`, `pom.xml`\n- 不生成密集小字表格\n\nDesign tokens:\n- background: #F7F5F0\n- surface: #FFFFFF\n- surface_tint_blue: #EEF4FB\n- surface_tint_teal: #EEF7F5\n- surface_tint_amber: #FCF6EA\n- primary: #1C2E4A\n- secondary: #2C7A7B\n- accent: #B7791F\n- text_primary: #1F2937\n- text_secondary: #667085\n- border: #D8DEE8\n- line_main: 2px\n- line_aux: 1px\n- radius_card: 10px\n- radius_pill: 999px\n\nVisual hierarchy:\n- controller / orchestrator / core object 权重最高\n- main path 最连续、最醒目\n- output / terminal 必须有明确收束感\n- decision / gate 必须像关键判断点\n- storage 稳定低调\n- observer / tracing 使用低对比虚线\n- tools 像可调用能力，不与主流程平级\n- annotation 收纳为侧栏、底栏或微型说明\n\nNode form rules:\n- event = compact pill\n- process = calm rectangle\n- resolver = compact structured block\n- decision / gate = logic block or split gate\n- artifact = distinct anchor object\n- output = terminal capsule\n- storage = grounded subtle container\n- reference = quiet chip\n- observer = low-contrast strip\n不要所有节点同尺寸，不要所有节点都画成白色矩形卡片。\n\nConnection rules:\n- Sequential flow: strongest line\n- Branch / merge: secondary line\n- Loop: curved return line\n- Handoff: highlighted but restrained connector\n- Lookup / reference: thin line\n- Observation: low-contrast dashed line\n- Dependency: desaturated structural line\n- Bidirectional exchange: two-way connector\n- Traceability: fine dashed line\n主流程线最清楚，辅助线降噪。\n箭头头部小而精确。\n禁止所有线同色同粗同风格。\n\nTypography:\n- modern clean technical editorial sans-serif feeling\n- title restrained, not poster-like\n- subtitle quieter than title\n- section heading clear\n- node title readable and prioritized\n- supporting text secondary\n- edge labels minimal but readable\n- Chinese and English mixed typesetting must be aligned and professional\n- code token should look monospace-like\n- avoid excessive bold text\n\nIcon rules:\n- icons are optional and secondary\n- icon area must not exceed 12% of node area\n- icon size must be smaller than 1.2× node title height\n- use consistent small line icons only\n- no cartoon icons\n- no oversized decorative icons\n- no icon on every node\n- structure and typography must carry more weight than icons\n\nBackground and finish:\n- warm off-white or soft technical canvas\n- very subtle paper grain or micro-grid allowed\n- extremely low visibility only\n- no heavy shadow\n- no glow\n- no glassmorphism\n- no 3D\n- no loud gradients\n- no colorful poster feeling\n\nLegend:\n- avoid legend if possible\n- if needed, make it extremely small and low contrast\n- legend must occupy less than 5% of image height\n\nLanguage rule:\n- node labels follow input language\n- code tokens remain in English\n- annotations match input language unless user specifies otherwise\n\nOverload handling:\n- if total entities &gt; 30, group aggressively into &lt;= 6 visible clusters\n- if the image source is blurry or partial, only keep confirmed information\n- if semantics are incomplete, omit uncertain nodes rather than inventing them\n\nUnsupported / irregular input fallback:\n如果结构无法归入常见类型，退化为：\n- one clear main mechanism\n- one visible primary path\n- clustered secondary relations\n- minimal annotation panel\n\nGood visual pattern:\n- one clear visual anchor\n- main mechanism dominates the center of the layout\n- auxiliary information is quieter and placed outside the main reading corridor\n- icons are small and restrained\n- spacing is generous\n- section boundaries rely on light surfaces and whitespace\n- the image feels like a premium technical editorial infographic\n\nBad visual pattern:\n- large decorative icons\n- all nodes same size\n- thick colorful cards everywhere\n- rainbow colors\n- heavy bottom legend bar\n- oversized title\n- strong gradients or glow\n- looks like a beautified Mermaid\n- looks like a generic PPT flowchart\n\nHard bans:\n- 不要漂亮版 Mermaid\n- 不要普通流程图换皮\n- 不要原样保留 subgraph 大框\n- 不要每个节点加圆形图标\n- 不要图标喧宾夺主\n- 不要彩虹配色\n- 不要大面积高饱和色块\n- 不要重阴影\n- 不要发光\n- 不要 3D\n- 不要玻璃拟态\n- 不要为了完整塞满所有文字\n- 不要编造不存在的信息\n\nSelf-check before rendering:\n1. 语义是否守恒\n2. 是否没有编造信息\n3. 主机制是否一眼可见\n4. 主路径是否清楚\n5. 布局是否明显不同于原 Mermaid / 原图\n6. 节点角色是否通过形态与权重区分\n7. 图标是否克制\n8. 色彩是否低饱和且高级\n9. 背景是否精致但不抢眼\n10. 文字是否清晰可读\n11. 线条是否有层级\n12. 画面是否避免了 AI 粗糙感和 PPT 模板感\n\n若任一项不合格，先重构画面，再输出最终信息图。\n\n输入 Mermaid Code 或图片：\n{粘贴 Mermaid / C4 / Flowchart / Sequence / State / ER / Timeline 源码，或上传渲染图片}\n</code></pre>\n<h2>什么时候用</h2>\n<p>如果要画架构总览、机制拆解、流程说明、复盘结论，优先看这一篇。输入内容越克制，效果越稳。复杂系统不要指望一张图讲完，先把主链路和关键判断喂给模型，剩下的细节留给正文。</p>\n","date_published":"2026-05-23T00:00:00.000Z","tags":["AI","提示词","生图","gpt-image-2","设计"],"language":"zh"},{"id":"https://www.lihuanyu.com/posts/gpt-image-2-cultural-poster-prompts/","url":"https://www.lihuanyu.com/posts/gpt-image-2-cultural-poster-prompts/","title":"gpt-image-2 提示词分享：诗词文化与东方叙事海报","summary":"整理适合诗词鉴赏、文化主题、东方幻想人像、章节页和叙事海报的 gpt-image-2 提示词模板。","content_html":"<p>诗词和文化类图片很容易滑向两个方向：要么只剩风景，要么只剩文字说明。前者像壁纸，后者像课件。</p>\n<p>这一篇放的是更偏文化表达和叙事海报的 prompt。它们会让模型先理解意象，再决定画面结构、文字位置、人物动作和空间层次。</p>\n<p>这个系列里的模板主要来自 @xiaoxiaodong01 和 @MrLarus 的公开分享，在这里一并致谢。这里记录的是我用 <code>gpt-image-2</code> 跑过以后，觉得还值得复用的用法。</p>\n<p>系列入口见 <a href=\"/posts/gpt-image-2-prompt-gallery/\">gpt-image-2 优秀提示词分享：可复用的生图模板</a>。</p>\n<h2>案例 1：纪念碑谷气质海报</h2>\n<p>同一句“太阳就站在黑夜的门口”，换一个 prompt，味道就完全变了。</p>\n<p>前一个案例靠电子元件建立微缩世界，这个案例靠建筑和几何空间。它的关键约束很清楚：不要把中文笔画直接做成立体字，而是先理解语义，再转成门洞、台阶、桥梁、平台和极小的人。</p>\n<p>这类图适合做章节页、封面图和带一点诗意的社媒配图。它不负责讲清复杂信息，只负责给一个概念留下一张安静、干净、能记住的画面。</p>\n<p><img src=\"/assets/posts/2026/gpt-image2-prompts/04-monument-valley-poster.jpg\" alt=\"纪念碑谷气质海报示例：太阳就站在黑夜的门口\"></p>\n<p>生成备注：</p>\n<ul>\n<li>模型：<code>gpt-image-2</code></li>\n<li>尺寸：<code>1024x1024</code></li>\n<li>输出：JPEG，压缩后约 58 KB</li>\n<li>测试输入：太阳就站在黑夜的门口</li>\n<li>适合：艺术海报、PPT 章节页、封面图、社媒配图</li>\n</ul>\n<p>可复制模板：</p>\n<pre><code class=\"language-text\">请根据用户最后输入的【主题 / 单词 / 短句】，生成一张「纪念碑谷气质」的极简超现实主义 3D 艺术海报。\n\n核心方式：\n先理解输入内容的含义、情绪与意象，再将其转译为一组具有建筑感的几何空间，而不是机械地把中文笔画直接做成立体字。若输入为中文，可提炼其核心语义，将关键词、氛围或象征元素巧妙融入主体造型、空间结构与排版中。\n\n画面主体：\n以主题语义生成纪念碑谷式的立体几何场景，形成平台、墙体、门洞、通道、桥梁、台阶、迷宫或悬浮结构。整体采用轻俯视等距视角 isometric / axonometric，画面像高级设计海报、建筑概念模型或艺术杂志封面。\n\n色彩与质感：\n整体保持高级、明亮、干净、适合印刷输出。主色以奶白、浅灰、砂岩、浅雾粉、淡蓝、薄荷绿、浅杏、柔和暖灰等低饱和浅色为主，减少大面积暗部。仅根据主题语义加入少量点亮色，精准用于局部边缘、深处空间、切面、门洞深处或微小装饰细节。材质为哑光微水泥、石膏、磨砂树脂或高级纸雕质感，光线柔和通透，层次清晰。\n\n构图与层次：\n主体仍为视觉中心，但四周需加入少量与主体同语义、同造型逻辑的呼应装饰，形成完整海报感，避免周围过空。装饰必须克制、简洁，并自然形成前景 / 中景 / 后景或主体 / 辅助 / 微装饰的层次关系，让画面更生动但不杂乱。可加入少量重复几何、悬浮小体块、细线结构、局部符号、边角呼应元素或微型景观。\n\n氛围元素：\n加入一到两个极简意境元素即可，例如微光、淡月、细枝、小鸟剪影、轻雾、花瓣、远处小体块或象征性图形，使画面更有诗意，但不要堆砌。\n\n人物尺度：\n可加入一到两个极小比例人物，作为情绪锚点，人物不超过画面高度 8%，动作安静自然，如站立、行走、眺望、穿行，不喧宾夺主。\n\n排版要求：\n可加入极少量设计感文字，如小标题、编号、年份、vol.01、竖排辅助文字或一句简短英文手写句。排版需服务画面，不可过多。\n\n最终要求：\n整体简洁、高级、轻盈、富有巧妙层次，兼具纪念碑谷式空间感、语义转译能力、海报设计感与印刷级配色控制。\n\n用户输入主题：\n{请输入主题、单词或短句}\n</code></pre>\n<h2>案例 2：印象派信息图海报</h2>\n<p>诗词类图片有个常见问题：画面挺美，但像壁纸，不像一页能拿去展示的内容。</p>\n<p>所以我用“小荷才露尖尖角，早有蜻蜓立上头”测这条。它要求插画占大部分画面，同时把文字主要放在左侧，分成顶部、中段、底部三个信息层级。</p>\n<p>成图里中文主标题可读，左侧信息图的层次也基本出来了。小字仍然要人工检查，尤其是模型自己补的解释，不能直接拿去给学生当知识点。</p>\n<p><img src=\"/assets/posts/2026/gpt-image2-prompts/06-impressionist-poster.jpg\" alt=\"印象派信息图海报示例：小荷才露尖尖角\"></p>\n<p>生成备注：</p>\n<ul>\n<li>模型：<code>gpt-image-2</code></li>\n<li>尺寸：<code>1536x1024</code></li>\n<li>输出：JPEG，压缩后约 264 KB</li>\n<li>测试输入：小荷才露尖尖角，早有蜻蜓立上头</li>\n<li>适合：PPT 首页、语文课件、诗词封面、编辑型海报</li>\n</ul>\n<p>可复制模板：</p>\n<pre><code class=\"language-text\">请根据用户输入的【主题 / 关键词 / 短句 / 诗句】，生成一张兼具艺术气质与信息图排版感的编辑型海报。\n\n整体风格为：平面插画、白色文字背景、印象派与后印象派气质的笔触、近景、线条抽象简约、柔焦效果、空气感、光斑笔刷，整体氛围梦幻、空灵、明亮、鲜艳但不脏乱。插画元素应覆盖画面的大部分区域，形成强烈的整体美感，色彩根据主题灵活变化，但始终保持高级、清透、柔和、有呼吸感；若主题适合柔美花意，可偏粉紫、浅粉、雾紫、嫩绿，但不要把粉紫写死，应让色彩从主题中自然生长。\n\n这不是普通配图海报，也不是满页杂乱文字，更不是左右均分版式。请保留“信息图海报”的排版逻辑：画面主体由满版或大面积插画构成，文字主要集中在左侧，并明确分布为三个层级区域，而不是缩成一坨。\n\n第一，左侧顶部必须形成一个清晰的信息簇，包含小字眉题、年份、英文小标题、中文主标题或副标题等，整体左对齐，字号有大有小，彼此拉开距离，形成上方的视觉起点。\n第二，左侧中段或左侧侧边，分布若干组小型信息内容或知识点，可做成短段落、边注、提示语、微型说明、关键词组等，数量约 4–6 组，要求高低错落、疏密变化、长短不一，像信息图海报中的辅助阅读区域，不能整齐堆成同一种文本块。\n第三，左侧底部要有收束性的文字区，可放一句中文短句、诗意总结、引导语，以及一小段英文说明或极小字注释，形成版面下方的落点，同样保持左对齐，并与顶部、中段形成明显区分。\n\n所有文字都必须像“被设计过的信息图排版”，具有上下呼应、大小变化、位置分层和节奏感：左上、左中、左下彼此独立又相互联系。不要把所有文字挤在同一块区域，不要平均排布，不要做成单一文本栏，不要缩成一团。\n\n插画不是装饰角料，而应成为主要视觉主体，占据画面大部分空间；文字则作为左侧的信息图式装饰与阅读结构存在。图像负责氛围与审美，文字负责信息层次与阅读趣味，整体像一张高级艺术信息海报：先被画面吸引，再被左侧版式打动，最后愿意停下来阅读。\n\n现在请围绕以下输入生成完整海报：\n\n【用户输入：{请输入主题、关键词、短句或诗句}】\n\n比例：16:9\n</code></pre>\n<h2>案例 3：几何情绪窗口诗词海报</h2>\n<p>诗词鉴赏图最怕两头不靠。</p>\n<p>只做风景，就成了壁纸；只做知识点，又容易像课件。这个 prompt 的解法是把真实物象放进一个窄长的几何色块里，让荷叶、花苞、水面这些具体东西从色块里长出来，再用留白承接文字。</p>\n<p>我用“接天莲叶无穷碧”测试。成图里的竖排标题、荷叶破框和右侧四个鉴赏点都比较清楚，也保住了白底、低饱和、日系摄影感这些要求。小字依然要人工检查，但它已经不是一张单纯的荷塘照片，而是能承担诗词鉴赏任务的海报。</p>\n<p><img src=\"/assets/posts/2026/gpt-image2-prompts/10-geometric-window-poster.jpg\" alt=\"几何情绪窗口诗词海报示例：接天莲叶无穷碧\"></p>\n<p>生成备注：</p>\n<ul>\n<li>模型：<code>gpt-image-2</code></li>\n<li>尺寸：<code>1536x1024</code></li>\n<li>输出：JPEG，压缩后约 124 KB</li>\n<li>测试输入：接天莲叶无穷碧</li>\n<li>适合：诗词鉴赏、节气海报、文化主题封面、美学杂志式配图</li>\n</ul>\n<p>可复制模板：</p>\n<pre><code class=\"language-text\">设计一张具有高级商业审美的极简海报，核心视觉语言是“真实物象穿越几何情绪窗口”。画面中设置一个窄长的低饱和色块，作为视觉锚点和空间容器，色块颜色根据主题选择柔和浅色，如雾蓝、浅青、米白、淡粉、暖灰或浅金。将【核心物象】以真实摄影质感或精细写实方式置入色块之中，但不要完全困在色块内，要让主体局部越界、破框、延伸到留白区域，形成自然生长感和空间穿透感。\n\n背景保持极简，使用大面积白色或浅灰留白，加入几乎透明的文化纹样、线性图形、地形线、水波线、光影轮廓或抽象符号，作为若隐若现的视觉细节。整体排版要像高端地产、奢侈品、美学杂志或节气海报，文字细长、克制、字距舒展，标题可竖排，辅助信息用小字号规整排列。画面需要有东方留白、现代秩序、自然生命力、轻奢品质感。\n\n避免杂乱、避免高饱和、避免厚重阴影、避免廉价模板感。\n\n本次主题：{请输入主题}\n主视觉元素：{请输入摄影风格、光影、质感要求}\n核心物象：{请输入主体物象}\n其他要求：画面至少 4 个不同层级、不同逻辑和不同呈现方案的知识点；知识点要服务主题，不要堆满。\n\n用途：{请输入用途}\n背景：{请输入背景要求}\n比例：{请输入比例，例如 16:9}\n</code></pre>\n<h2>案例 4：东方奇幻叙事人像海报</h2>\n<p>人物海报最容易被模型画成两类东西：游戏立绘，或者古风写真。</p>\n<p>这条 prompt 的核心约束不是“好看的古风美人”，而是“画中人破卷而出”。它把画卷当成中轴结构，让人物的一部分留在画里，脸、手和上半身冲到画外。这样画面就有了叙事关系，而不只是一个站姿漂亮的人。</p>\n<p>我测试时填的是“月下仙子”。成图里长画卷足够明显，人物伸手朝向镜头，裙摆和背景还留在画卷之内，前景白梅、飘纱、月亮和水面也撑起了空间层次。手部细节仍然是人物图最需要检查的地方，但整体破界感已经出来了。</p>\n<p><img src=\"/assets/posts/2026/gpt-image2-prompts/11-scroll-beauty-poster.jpg\" alt=\"东方奇幻叙事人像海报示例：画中美人破卷而出\"></p>\n<p>生成备注：</p>\n<ul>\n<li>模型：<code>gpt-image-2</code></li>\n<li>尺寸：<code>1024x1792</code></li>\n<li>输出：JPEG，压缩后约 217 KB</li>\n<li>测试输入：月下仙子</li>\n<li>适合：手机壁纸、古风人像海报、角色视觉、竖版封面、东方幻想主题图</li>\n</ul>\n<p>可复制模板：</p>\n<pre><code class=\"language-text\">请生成一张高完成度的手机竖版 9:16 东方奇幻叙事人像海报，系列名称为「画中美人 · 破卷而出」。\n\n【主题设定】\n主题人物：【填写，如：古风美人 / 花神 / 月下仙子 / 冷月神女 / 红衣剑姬 / 鹤影仙姬】\n人物气质：【填写，如：楚楚可怜 / 温柔大方 / 清冷疏离 / 明艳凌厉 / 妖冶神秘 / 庄重神圣】\n表情情绪：【填写，如：眼神湿润、楚楚动人 / 温柔含笑 / 冷淡克制 / 锐利坚定 / 哀伤欲言又止】\n动作设定：【填写，如：向前伸手 / 半步迈出画卷 / 身体前倾探出 / 飞身破卷而出 / 扶袖回眸】\n服装方向：【填写，如：白青轻纱汉服 / 粉金温柔长裙 / 冰蓝月色仙裙 / 红金华丽古风衣袍】\n主色调：【填写，如：白+淡青+浅金 / 米白+柔粉+杏金 / 青白+冰蓝+银灰 / 赤红+金+奶白】\n花材元素：【填写，如：白梅 / 玉兰 / 海棠 / 牡丹 / 山茶 / 樱花】\n辅助元素：【填写，如：仙鹤 / 飞鸟 / 飘纱 / 花瓣 / 水面倒影 / 雾气 / 金色微粒 / 月亮 / 楼阁】\n\n【核心视觉骨架】\n画面中心必须有一幅非常明显的纵向长画卷 / 挂轴，作为整个画面的核心结构。人物不是站在画卷前，而是必须呈现出明显的“从画卷中探出来 / 走出来 / 飞出来 / 破卷而出”的视觉效果。\n\n必须同时满足以下结构逻辑：\n1. 人物的大半个身体已经离开画卷，整体趋势明显朝向镜头前方，有很强的破界感和前冲感。\n2. 必须保留一部分身体、裙摆、衣摆、脚步或发丝仍然留在画卷之内，形成清晰的“画里画外融合”。\n3. 画中人物部分可略带柔和、朦胧、绘画感；探出画外的脸部、手部、上半身则要更加写实、清晰、细腻。\n4. 人物与画卷、花枝、飞鸟、飘纱要形成穿插关系，而不是平面摆拍。\n\n【画卷内容】\n画卷必须足够长，也比普通卷轴略宽，整体舒展、大气、完整。画卷之内可呈现与主题匹配的东方画境，如：\n- 山水云雾\n- 花树花境\n- 月夜楼阁\n- 水墨亭台\n- 抽象东方秘境\n- 鹤影寒江\n- 花神卷境\n\n【人物表现】\n人物必须是高颜值、精致、写实的东方人物形象，具有电影级写实人像质感。人物表情要有情绪感染力，不能呆板，不能恐怖，不能女鬼化。整体感觉应唯美、心动、有故事感，像“画中人来到现实”的惊艳瞬间。\n\n【前景与空间层次】\n画面必须有明显但不喧宾夺主的前景元素，用于制造真实空间纵深，例如：\n- 前景花枝 / 花朵\n- 少量仙鹤或飞鸟\n- 飘动轻纱\n- 飘落花瓣\n- 水面倒影\n- 柔和雾气\n- 金色微粒 / 光点\n\n注意：\n1. 前景元素要像真实存在于物理空间中，而不是只贴在背景上。\n2. 仙鹤 / 飞鸟只能作为陪衬点缀，不能比人物更抢眼。\n3. 前景元素要帮助建立“现实空间”，从而强化人物已经破卷而出的感觉。\n\n【背景要求】\n背景不要纯黑，也不要复杂写实大场景。请使用低信息量暗色氛围背景，如：\n- 深青灰雾境\n- 月夜柔雾花境\n- 水墨云烟背景\n- 暖金暗雾背景\n\n背景需要有层次感、空气感、轻微光晕与雾感，但整体仍应克制，不能喧宾夺主。目的是营造“画中世界与现实空间轻度重叠”的梦幻舞台氛围。\n\n【构图与镜头感】\n整体构图为 9:16 竖版海报式构图，以中轴长画卷为稳定核心，以人物面部和伸出的手部为视觉焦点。上部保留适当呼吸空间和氛围元素，中部强化人物表情与动作，下部通过衣摆、水面、花枝完成落地与收束。整体要有明确阅读路径：先看脸，再看手，再理解破卷而出的叙事关系。\n\n【整体风格要求】\n整体风格为：\n东方幻想美学、古风电影感、写实唯美人像、装置式花境、叙事型海报、梦幻但清晰、精致高完成度。\n\n强调：\n- 画面清晰，不要过度模糊\n- 探出画外的部分更写实\n- 构图有强烈纵深和镜头感\n- 质感高级，不要低质插画感，不要普通古风写真模板感，不要游戏立绘感\n- 最终效果要像一张高级东方幻想视觉海报\n\n【如果用户上传人物参考图】\n如果有参考人物图，请以 Image A 作为人物身份参考，保留其五官辨识度、脸型、妆感方向和整体气质，在此基础上完成「画中美人 · 破卷而出」风格化创作，不要换成陌生脸。\n\n请输出一张细节丰富、氛围唯美、具有明显“从画卷中破卷而出”视觉冲击力的最终成品图。\n</code></pre>\n<h2>什么时候用</h2>\n<p>如果要做诗词鉴赏页、节气海报、章节页、古风人像或文化主题封面，可以从这里找。人物类图尤其要多看一眼手、眼神和肢体关系；诗词类图则要检查标题和知识点，不要让模型自己把诗意解释歪。</p>\n","date_published":"2026-05-23T00:00:00.000Z","tags":["AI","提示词","生图","gpt-image-2","设计"],"language":"zh"},{"id":"https://www.lihuanyu.com/posts/gpt-image-2-prompt-gallery/","url":"https://www.lihuanyu.com/posts/gpt-image-2-prompt-gallery/","title":"gpt-image-2 提示词库：技术图、商业封面与视觉海报模板","summary":"按用途整理经过实测的 gpt-image-2 提示词，覆盖技术图解、信息图、商业封面、微缩视觉、文化海报和城市壁纸。","content_html":"<p>最近一边写博客，一边给技术文章和分享页补图，我越来越觉得，生图 prompt 不该写成咒语。</p>\n<p>一句“生成一张高级感科技海报”当然也能跑。运气好的时候，模型会给一张还不错的图；运气差的时候，就是大字发糊、信息乱飞、构图像从模板库里随机拼出来的封面。</p>\n<p>很多时候，模型不是不会画，是需求没有被说清楚。</p>\n<p>比较好用的 prompt，更像一份设计 brief。它会告诉模型：这张图要干什么，谁是主角，信息怎么取舍，文字怎么分层，什么东西不能乱编，最后用户只需要替换哪一块输入。</p>\n<p>下面这些模板主要来自 @xiaoxiaodong01 和 @MrLarus 的公开分享，在这里一并致谢。我把它们用 <code>gpt-image-2</code> 跑了一遍，把能复用的场景、实测图和原始 prompt 放在一起。单篇文章塞太多长 prompt 会很难读，所以按用途拆成几组。后面继续补，也不会把一篇文章撑成瀑布。</p>\n<p><a href=\"/en/posts/2026/gpt-image-2-prompt-patterns/\">English version: Reusable gpt-image-2 Prompt Patterns</a></p>\n<p>工具不重要。可以用我前面写过的 <a href=\"/posts/2026/%E4%BB%8E-tinify-cli-%E5%88%B0-imgasset-%E6%8A%8A%E5%8D%9A%E5%AE%A2%E9%85%8D%E5%9B%BE%E5%81%9A%E6%88%90%E4%B8%80%E6%9D%A1%E6%B5%81%E6%B0%B4%E7%BA%BF/\">从 tinify-cli 到 imgasset：把博客配图做成一条流水线</a>，也可以直接把 prompt 粘到图片生成工具里。关键是把输入、输出和最终用途留下来。只收藏一段 prompt，不知道它在哪个主题上跑过，过一阵基本也就忘了。</p>\n<p>先说几个不太浪漫的前提。</p>\n<p>第一，模型名、价格、可用尺寸和平台入口会变。长期使用时以官方页面为准，文章里更值得留下的是 prompt 的结构。</p>\n<p>第二，图片模型生成文字仍然要人工校对。标题、大字、短标签通常比较稳，密集小字、长句、专有名词和数字最容易出错。</p>\n<p>第三，涉及年份、市场变化、政策、价格、榜单、人物、机构、品牌、新闻等会变化的信息，不要让图片模型自己补。更稳的做法是先把可靠资料整理成输入，再让模型做视觉转译。</p>\n<h2>先按用途选</h2>\n<div class=\"table-scroll\"><table>\n<thead>\n<tr>\n<th>想做什么</th>\n<th>从这里开始</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>架构图、机制拆解、Mermaid 流程重构</td>\n<td><a href=\"/posts/gpt-image-2-technical-diagram-prompts/\">技术图解与信息图</a></td>\n</tr>\n<tr>\n<td>PPT 封面、报告插图、知识卡片</td>\n<td><a href=\"/posts/gpt-image-2-commercial-cover-prompts/\">商业封面与微缩视觉</a></td>\n</tr>\n<tr>\n<td>诗词、节气、东方幻想人像</td>\n<td><a href=\"/posts/gpt-image-2-cultural-poster-prompts/\">诗词文化与东方叙事海报</a></td>\n</tr>\n<tr>\n<td>城市旅行海报、桌面或手机壁纸</td>\n<td><a href=\"/posts/gpt-image-2-city-wallpaper-prompts/\">城市海报与艺术壁纸</a></td>\n</tr>\n</tbody>\n</table>\n</div><p>每一组都保留了实测输入、生成结果、完整 prompt 和适用边界。这个总入口只负责帮人找到正确模板，不重复堆放长 prompt。</p>\n<h2>我为什么收这些 prompt</h2>\n<p>我挑 prompt 时，不看它写得长不长。长 prompt 很多，真正能反复用的不多。</p>\n<p>我主要看几件事。</p>\n<ul>\n<li>场景明确：知道它适合封面、海报、报告配图、技术图、壁纸，还是社媒头图。</li>\n<li>信息有边界：不会要求模型把所有东西都塞进一张图，而是主动压缩重点。</li>\n<li>视觉有秩序：标题、主体、辅助信息、留白、动线和收束点都有安排。</li>\n<li>失败有约束：明确禁止乱码、虚假数据、过度拥挤、廉价模板感、错误组件等常见问题。</li>\n<li>输入可替换：最后只要改主题、短句、资料或 Mermaid 源码，就能复用。</li>\n</ul>\n<p>放在一起看，反而能看出一个规律：好 prompt 不能只堆形容词，它得替模型安排工作。</p>\n<h2>系列目录</h2>\n<h3>技术图解与信息图</h3>\n<p>适合技术方案、架构说明、机制拆解、流程说明和 Mermaid 重构。</p>\n<p><img src=\"/assets/posts/2026/gpt-image2-prompts/01-handdrawn-knowledge.jpg\" alt=\"技术图解与信息图示例\"></p>\n<p>完整模板见 <a href=\"/posts/gpt-image-2-technical-diagram-prompts/\">gpt-image-2 提示词分享：技术图解与信息图</a>。</p>\n<h3>商业封面与微缩视觉</h3>\n<p>适合 PPT 封面、报告插图、自媒体头图、科技行业封面、微缩视觉、品牌 Logo 概念稿和知识卡片。</p>\n<p><img src=\"/assets/posts/2026/gpt-image2-prompts/13-blue-plumbago-knowledge-card.jpg\" alt=\"商业封面与微缩视觉示例\"></p>\n<p>完整模板见 <a href=\"/posts/gpt-image-2-commercial-cover-prompts/\">gpt-image-2 提示词分享：商业封面与微缩视觉</a>。</p>\n<h3>诗词文化与东方叙事海报</h3>\n<p>适合诗词鉴赏、节气海报、章节页、东方幻想人像和文化主题封面。</p>\n<p><img src=\"/assets/posts/2026/gpt-image2-prompts/11-scroll-beauty-poster.jpg\" alt=\"诗词文化与东方叙事海报示例\"></p>\n<p>完整模板见 <a href=\"/posts/gpt-image-2-cultural-poster-prompts/\">gpt-image-2 提示词分享：诗词文化与东方叙事海报</a>。</p>\n<h3>城市海报与艺术壁纸</h3>\n<p>适合桌面壁纸、手机壁纸、城市旅行海报、拼音主题海报和收藏视觉。</p>\n<p><img src=\"/assets/posts/2026/gpt-image2-prompts/09-hangzhou-romanized-poster.jpg\" alt=\"城市海报与艺术壁纸示例\"></p>\n<p>完整模板见 <a href=\"/posts/gpt-image-2-city-wallpaper-prompts/\">gpt-image-2 提示词分享：城市海报与艺术壁纸</a>。</p>\n<h2>让 Agent 直接跑这些 prompt</h2>\n<p>我现在更倾向于把生图也当成一次小型构建。</p>\n<p>提示词模板、用户输入、原图、压缩图、最终引用，最好都能留下记录。这样一篇文章、一个系列封面、一组 PPT 图，后面要重跑、换主题、换尺寸，才知道从哪里开始。</p>\n<p>以 <code>imgasset</code> 为例，可以写一个 JSONL 文件：</p>\n<pre><code class=\"language-jsonl\">{&quot;out&quot;:&quot;prompt-share/demo.png&quot;,&quot;size&quot;:&quot;1536x1024&quot;,&quot;prompt&quot;:&quot;请把下面这段中文提示词原样作为 gpt-image-2 的 prompt，不要翻译、不要改写、不要解释。\\n\\n{把完整提示词粘贴在这里}&quot;}\n</code></pre>\n<p>然后生成原图：</p>\n<pre><code class=\"language-bash\">imgasset generate temp/gpt-image2-prompt-test/prompts.jsonl \\\n  --raw-dir temp/gpt-image2-prompt-test/raw \\\n  --skip-existing\n</code></pre>\n<p>再把选中的图压缩发布到站点资产目录：</p>\n<pre><code class=\"language-bash\">tinify temp/gpt-image2-prompt-test/raw \\\n  --recursive \\\n  --out-dir public/assets/posts/2026/gpt-image2-prompts \\\n  --format jpeg \\\n  --background white \\\n  --suffix &quot;&quot;\n</code></pre>\n<p>没有命令行链路也没关系，直接把模板粘到图片生成工具里也能用。工具形态没那么重要，别只留下最后那张图就行。只留图，不留输入，下次想复现就只能靠猜。</p>\n<h2>最后</h2>\n<p>提示词越长，不一定越好。</p>\n<p>但如果一段 prompt 能稳定表达任务、边界、层级、风格和禁区，它就比一句风格词可靠很多。</p>\n<p>这也是我继续收集这些 prompt 的原因。真正有价值的不是某一张图，而是一套能反复替换主题、反复验证、反复用于内容生产的视觉工作流。</p>\n","date_published":"2026-05-23T00:00:00.000Z","date_modified":"2026-07-31T00:00:00.000Z","tags":["AI","提示词","生图","gpt-image-2","设计"],"language":"zh"},{"id":"https://www.lihuanyu.com/en/posts/2026/from-tinify-cli-to-imgasset-turning-blog-images-into-a-pipeline/","url":"https://www.lihuanyu.com/en/posts/2026/from-tinify-cli-to-imgasset-turning-blog-images-into-a-pipeline/","title":"From tinify-cli to imgasset: Turning Blog Images into a Pipeline","summary":"How a repeated blog image workflow became two reusable npm CLI tools: tinify-cli for compression and imgasset for AI generation, raw asset storage, compression, and publishing.","content_html":"<p>While adding images to recent blog posts, I ran into a familiar kind of work: the process itself was not difficult, but it had too many small steps, and the steps were almost the same every time.</p>\n<p>First, read the article and write prompts. Then generate raw images with an image model. The originals should not go into the repository, so they need to stay in a temporary directory. The selected images need to be compressed, converted into a web-friendly format, and placed under <code>public/assets/</code>. Finally, the Markdown image paths need to be added, followed by a build and link check.</p>\n<p>Doing this manually for one or two posts is fine. After enough posts, it becomes repetitive work.</p>\n<p>So I split the workflow into two tools:</p>\n<ul>\n<li><a href=\"https://github.com/YGM-Studio/tinify-cli\"><code>@yigemo/tinify-cli</code></a>: image compression and format conversion.</li>\n<li><a href=\"https://github.com/YGM-Studio/imgasset\"><code>@yigemo/imgasset</code></a>: AI image generation, raw image storage, compression, and publishing output.</li>\n</ul>\n<p>They were not designed as polished products from day one. They came out of the real process of writing and maintaining this blog, one layer at a time.</p>\n<p><a href=\"/posts/2026/%E4%BB%8E-tinify-cli-%E5%88%B0-imgasset-%E6%8A%8A%E5%8D%9A%E5%AE%A2%E9%85%8D%E5%9B%BE%E5%81%9A%E6%88%90%E4%B8%80%E6%9D%A1%E6%B5%81%E6%B0%B4%E7%BA%BF/\">Chinese version of this article</a></p>\n<p><img src=\"/assets/posts/2026/image-asset-pipeline/01-image-pipeline.jpg\" alt=\"AI image generation, compression, and publishing directories organized as one image asset pipeline\"></p>\n<p>To run the full workflow, install <code>imgasset</code> first:</p>\n<pre><code class=\"language-bash\">npm install -g @yigemo/imgasset\n</code></pre>\n<p><code>imgasset</code> already includes <code>tinify-cli</code> as a dependency. If only image compression is needed, the compression tool can also be installed directly:</p>\n<pre><code class=\"language-bash\">npm install -g @yigemo/tinify-cli\n</code></pre>\n<h2>First Layer: Standardize Image Compression</h2>\n<p>Before writing <code>imgasset</code>, I first wrote <code>tinify-cli</code>.</p>\n<p>Image compression looks like a small problem, but it appears constantly on content sites. Blog images, official website illustrations, product screenshots, and social sharing images eventually face the same issue: raw files are too large to publish directly.</p>\n<p>TinyPNG / Tinify has always worked well for compression. But opening a website, uploading files, downloading results, or copying the same API script across projects is not a good long-term workflow.</p>\n<p>The goal of <code>tinify-cli</code> is simple: make compression a stable command.</p>\n<pre><code class=\"language-bash\">tinify login\n</code></pre>\n<p>Then compress a directory:</p>\n<pre><code class=\"language-bash\">tinify temp/article/raw \\\n  --recursive \\\n  --out-dir public/assets/article \\\n  --format jpeg \\\n  --background white \\\n  --suffix &quot;&quot;\n</code></pre>\n<p>Several options matter a lot for blog images.</p>\n<p><code>--recursive</code> preserves the directory structure, which is useful when one article has multiple images.</p>\n<p><code>--out-dir</code> writes compressed images into the publishing directory instead of overwriting originals.</p>\n<p><code>--format jpeg</code> and <code>--background white</code> convert PNG outputs into JPEG files that are usually more suitable for web pages, while handling transparent backgrounds.</p>\n<p><code>--suffix &quot;&quot;</code> keeps final file names clean. A raw file such as <code>temp/article/raw/01-context.png</code> can become <code>public/assets/article/01-context.jpg</code>.</p>\n<p><img src=\"/assets/posts/2026/image-asset-pipeline/02-compression-workbench.jpg\" alt=\"Raw images entering compression and format conversion before reaching the publishing directory\"></p>\n<p>This layer solves repeatable compression and format conversion.</p>\n<p>But another question appeared quickly: where should the raw images come from?</p>\n<h2>Second Layer: AI Image Generation Needs a Workflow Too</h2>\n<p>When adding images to articles, image generation is only one part of the job.</p>\n<p>The surrounding workflow is where the friction lives:</p>\n<ul>\n<li>Prompts should be saved and reused.</li>\n<li>Generated originals need a fixed location.</li>\n<li>Existing images should not be regenerated accidentally.</li>\n<li>API keys should never enter the project repository.</li>\n<li>Model, size, quality, proxy, and base URL settings should be reusable.</li>\n<li>After generation, images should be able to move directly into compression.</li>\n</ul>\n<p>If every project gets a temporary script, these details scatter quickly. After a while, the scripts become their own maintenance problem: which one still works, which one is outdated, which one hard-codes local paths, and which one might contain sensitive configuration.</p>\n<p><code>imgasset</code> handles this layer.</p>\n<p>It splits image generation configuration into three parts:</p>\n<ul>\n<li>Global profile: base URL, model, size, quality, proxy, and other non-sensitive settings.</li>\n<li>Global secret: API key, kept outside project repositories.</li>\n<li>Project configuration: raw directory, publishing directory, and compression options for the current project.</li>\n</ul>\n<p>Initialize configuration:</p>\n<pre><code class=\"language-bash\">imgasset config init\n</code></pre>\n<p>Create a profile:</p>\n<pre><code class=\"language-bash\">imgasset profile set default \\\n  --base-url https://api.example.com/v1 \\\n  --model gpt-image-2 \\\n  --size 1536x1024 \\\n  --quality medium \\\n  --output-format png \\\n  --default\n</code></pre>\n<p>Save the API key:</p>\n<pre><code class=\"language-bash\">imgasset secret set default\n</code></pre>\n<p>Then write a JSONL prompt file inside the project:</p>\n<pre><code class=\"language-jsonl\">{&quot;out&quot;:&quot;01-context.png&quot;,&quot;prompt&quot;:&quot;Minimal surreal isometric 3D editorial poster...&quot;}\n{&quot;out&quot;:&quot;02-flow.png&quot;,&quot;prompt&quot;:&quot;Minimal surreal isometric 3D editorial poster...&quot;}\n</code></pre>\n<p>Generate raw images:</p>\n<pre><code class=\"language-bash\">imgasset generate prompts.jsonl \\\n  --raw-dir temp/imgasset/article/raw \\\n  --skip-existing\n</code></pre>\n<p><code>--skip-existing</code> is especially useful. Image generation can be slow, and network failures are common enough to design for. If a batch fails on the third image, the next run should not regenerate the first two.</p>\n<h2>Connect the Two Steps</h2>\n<p>Using <code>tinify-cli</code> and <code>imgasset generate</code> separately already covers most cases. The smoother version is one command for the whole workflow:</p>\n<pre><code class=\"language-bash\">imgasset run prompts.jsonl \\\n  --raw-dir temp/imgasset/article/raw \\\n  --publish-dir public/assets/article \\\n  --format jpeg \\\n  --background white \\\n  --skip-existing\n</code></pre>\n<p>This command generates the originals first, then uses the bundled <code>tinify-cli</code> dependency for compression and format conversion.</p>\n<p><img src=\"/assets/posts/2026/image-asset-pipeline/03-generation-publish-flow.jpg\" alt=\"Prompts, raw image storage, and the publishing directory connected into one continuous production path\"></p>\n<p>In practice, installing <code>imgasset</code> gives a project both image generation and compression. There is no need to maintain a separate compression script in every repository.</p>\n<p>I usually keep the directory structure like this:</p>\n<pre><code class=\"language-text\">temp/\n  imgasset/\n    my-article/\n      raw/\npublic/\n  assets/\n    posts/\n      2026/\n        my-article/\nprompts.jsonl\n</code></pre>\n<p><code>temp/</code> holds originals and temporary files, and stays in <code>.gitignore</code>.</p>\n<p><code>public/assets/</code> holds compressed publishing assets that can be referenced by articles.</p>\n<p>Markdown only references the final output:</p>\n<pre><code class=\"language-markdown\">![Content system illustration](/assets/posts/2026/my-article/01-context.jpg)\n</code></pre>\n<p>That boundary matters. Originals are production material. Published images are website assets.</p>\n<h2>Why Not Just Use One Script</h2>\n<p>At the beginning, a single script is absolutely enough.</p>\n<p>For one project, it is often the fastest option. Read an API key from the environment, loop over image generation requests, then call the compression API. A few dozen lines of code can work.</p>\n<p>The problem is that this kind of script often becomes a disposable asset.</p>\n<p>When a second project needs the same workflow, the script gets copied. Then paths change, model names change, compression formats change, proxy settings change, and error handling changes. After a few rounds, it becomes hard to know which copy represents the current practice.</p>\n<p>The value of turning this into a tool is not fewer lines of code. The value is fixing the boundaries:</p>\n<ul>\n<li>API keys never enter projects.</li>\n<li>Originals default to temporary directories.</li>\n<li>Prompts are saved as JSONL.</li>\n<li>Output paths are controlled by commands or project config.</li>\n<li>Compression is provided as a dependency instead of a separate setup step.</li>\n<li>Interrupted runs can continue.</li>\n</ul>\n<p>Once these conventions are stable, the same workflow can move across projects.</p>\n<h2>Security and Open Source</h2>\n<p>Both tools are published to npm and available on GitHub.</p>\n<p>Before open sourcing them, I focused on three things.</p>\n<p>First, secrets. API keys should only live in global secret files or environment variables. Project configuration, prompt files, logs, and reports should not contain them.</p>\n<p><img src=\"/assets/posts/2026/image-asset-pipeline/04-config-secret-boundary.jpg\" alt=\"Configuration, secrets, temporary files, and public assets separated by clear boundaries\"></p>\n<p>Second, examples. Examples should use placeholders such as <code>https://api.example.com/v1</code>, not any real service endpoint from daily use.</p>\n<p>Third, publishing. Both packages follow a standard npm package shape. <code>imgasset</code> also uses GitHub Actions and npm Trusted Publishing. Releasing is done with:</p>\n<pre><code class=\"language-bash\">pnpm run release\n</code></pre>\n<p>The script increments the version, creates a tag, and GitHub Actions publishes to npm from that tag.</p>\n<p>This is a little more structured than running <code>npm publish</code> manually, but it is easier to trace over time. For open source packages, a clear release path is worth the extra setup.</p>\n<h2>It Still Comes Back to Writing</h2>\n<p>After finishing these tools, the biggest change was not that I type fewer commands. It was that adding images no longer interrupts the writing rhythm as much.</p>\n<p>Previously, the thought of adding images brought a list of small chores: where prompts should live, where originals should go, what the compressed files should be called, whether paths would be wrong, and how to resume after a failed run. Each detail was small, but together they pulled attention away from the article.</p>\n<p>Now the workflow feels closer to this:</p>\n<ol>\n<li>Read the article and decide how many images it needs.</li>\n<li>Write <code>prompts.jsonl</code>.</li>\n<li>Run <code>imgasset run</code>.</li>\n<li>Insert the output images into Markdown.</li>\n<li>Build and check.</li>\n</ol>\n<p>The real judgment is still manual: what imagery fits the article, how many images are enough, where an image helps reading, and which attractive image should still be rejected. The tools only fix the directories, commands, format conversion, and retry behavior.</p>\n<p>That is what this post is really about: not two npm packages, but a small process becoming stable.</p>\n<p>The best small tools are almost invisible in daily use, but immediately useful when moving to another project. <code>tinify-cli</code> handles compression. <code>imgasset</code> connects generation, raw storage, and publishing output. Together they solve not one image generation task, but a reusable image asset workflow for the next article, the next site, and the next project.</p>\n","date_published":"2026-05-17T00:00:00.000Z","tags":["Image Processing","AI","CLI","npm","Engineering"],"language":"en"},{"id":"https://www.lihuanyu.com/posts/2026/%E4%BB%8E-tinify-cli-%E5%88%B0-imgasset-%E6%8A%8A%E5%8D%9A%E5%AE%A2%E9%85%8D%E5%9B%BE%E5%81%9A%E6%88%90%E4%B8%80%E6%9D%A1%E6%B5%81%E6%B0%B4%E7%BA%BF/","url":"https://www.lihuanyu.com/posts/2026/%E4%BB%8E-tinify-cli-%E5%88%B0-imgasset-%E6%8A%8A%E5%8D%9A%E5%AE%A2%E9%85%8D%E5%9B%BE%E5%81%9A%E6%88%90%E4%B8%80%E6%9D%A1%E6%B5%81%E6%B0%B4%E7%BA%BF/","title":"从 tinify-cli 到 imgasset：把博客配图做成一条流水线","summary":"从图片压缩工具 tinify-cli 到 AI 生图与压缩一体化工具 imgasset，记录一套博客配图工作流如何从临时脚本沉淀成可复用的 npm CLI。","content_html":"<p>最近给博客补配图时，我又遇到了一个很熟悉的问题：流程本身并不复杂，但步骤很多，而且每次都差不多。</p>\n<p>先根据文章内容写提示词，用图片模型生成原图；原图不能直接进仓库，要放在临时目录里；选中的图要压缩、转成适合网页使用的格式，再放到 <code>public/assets/</code> 下；最后把 Markdown 里的图片路径补上，跑构建和链接检查。</p>\n<p>一两篇文章手工做没什么问题。文章多了之后，这件事就开始变成一种重复劳动。</p>\n<p>于是我把这个流程拆成了两个工具：</p>\n<ul>\n<li><a href=\"https://github.com/YGM-Studio/tinify-cli\"><code>@yigemo/tinify-cli</code></a>：负责图片压缩和格式转换。</li>\n<li><a href=\"https://github.com/YGM-Studio/imgasset\"><code>@yigemo/imgasset</code></a>：负责 AI 生图、原图保存、压缩和发布输出。</li>\n</ul>\n<p>它们不是一开始就设计好的“产品”。更准确地说，是从真实写博客的流程里，一层一层提炼出来的。</p>\n<p><a href=\"/en/posts/2026/from-tinify-cli-to-imgasset-turning-blog-images-into-a-pipeline/\">English version: From tinify-cli to imgasset: Turning Blog Images into a Pipeline</a></p>\n<p><img src=\"/assets/posts/2026/image-asset-pipeline/01-image-pipeline.jpg\" alt=\"AI 生图、压缩和发布目录被组织成一条图片资产流水线\"></p>\n<p>如果要跑完整流程，可以先安装 <code>imgasset</code>：</p>\n<pre><code class=\"language-bash\">npm install -g @yigemo/imgasset\n</code></pre>\n<p><code>imgasset</code> 已经把 <code>tinify-cli</code> 作为依赖带上了。只需要单独做图片压缩时，也可以只安装压缩工具：</p>\n<pre><code class=\"language-bash\">npm install -g @yigemo/tinify-cli\n</code></pre>\n<h2>第一层：先把图片压缩标准化</h2>\n<p>在写 <code>imgasset</code> 之前，我先写了 <code>tinify-cli</code>。</p>\n<p>图片压缩这件事看起来很小，但在内容站里非常高频。博客文章、官网配图、产品截图、社交分享图，最后都要面对同一个问题：原图太大，直接放线上不合适。</p>\n<p>TinyPNG / Tinify 的压缩效果一直不错，但如果每次都打开网页上传下载，或者在不同项目里复制一段调用 API 的脚本，长期维护会很麻烦。</p>\n<p>所以 <code>tinify-cli</code> 的目标很简单：把压缩变成一条稳定命令。</p>\n<pre><code class=\"language-bash\">tinify login\n</code></pre>\n<p>登录后压缩一个目录：</p>\n<pre><code class=\"language-bash\">tinify temp/article/raw \\\n  --recursive \\\n  --out-dir public/assets/article \\\n  --format jpeg \\\n  --background white \\\n  --suffix &quot;&quot;\n</code></pre>\n<p>这里几个参数对博客配图很关键。</p>\n<p><code>--recursive</code> 可以保留目录结构，适合一篇文章有多张图的情况。</p>\n<p><code>--out-dir</code> 可以把压缩后的图片写到发布目录，而不是覆盖原图。</p>\n<p><code>--format jpeg</code> 和 <code>--background white</code> 让生成图可以从 PNG 转成更适合网页展示的 JPEG，同时处理透明背景。</p>\n<p><code>--suffix &quot;&quot;</code> 则是为了让最终文件名保持干净。原图可能在 <code>temp/article/raw/01-context.png</code>，发布图可以直接变成 <code>public/assets/article/01-context.jpg</code>。</p>\n<p><img src=\"/assets/posts/2026/image-asset-pipeline/02-compression-workbench.jpg\" alt=\"原图经过压缩和格式转换后进入发布目录\"></p>\n<p>这一步解决的是压缩和格式转换的可重复性。</p>\n<p>但很快又出现了第二个问题：原图从哪里来？</p>\n<h2>第二层：AI 生图也需要工作流</h2>\n<p>给文章配图时，生图本身只是一个环节。</p>\n<p>真正麻烦的是围绕生图的上下文：</p>\n<ul>\n<li>每张图的提示词要能保存和复用。</li>\n<li>生成出来的原图要有固定位置。</li>\n<li>已经生成过的图不要重复生成。</li>\n<li>API key 不能写进项目仓库。</li>\n<li>模型、尺寸、质量、代理、base URL 这些配置应该可以复用。</li>\n<li>生成完之后最好能直接进入压缩流程。</li>\n</ul>\n<p>如果每次都临时写脚本，这些细节会不断散落在不同项目里。脚本一多，就会出现新的问题：哪个脚本能用，哪个脚本已经过期，哪个脚本里写死了本地路径，哪个脚本里不小心带了敏感配置。</p>\n<p><code>imgasset</code> 解决的就是这部分。</p>\n<p>它把生图配置分成三层：</p>\n<ul>\n<li>全局 profile：保存 base URL、模型、尺寸、质量、代理等非敏感配置。</li>\n<li>全局 secret：保存 API key，不进入项目仓库。</li>\n<li>项目配置：保存当前项目的原图目录、发布目录、压缩参数。</li>\n</ul>\n<p>初始化配置：</p>\n<pre><code class=\"language-bash\">imgasset config init\n</code></pre>\n<p>创建一个 profile：</p>\n<pre><code class=\"language-bash\">imgasset profile set default \\\n  --base-url https://api.example.com/v1 \\\n  --model gpt-image-2 \\\n  --size 1536x1024 \\\n  --quality medium \\\n  --output-format png \\\n  --default\n</code></pre>\n<p>保存 API key：</p>\n<pre><code class=\"language-bash\">imgasset secret set default\n</code></pre>\n<p>然后在项目里写一个 JSONL 提示词文件：</p>\n<pre><code class=\"language-jsonl\">{&quot;out&quot;:&quot;01-context.png&quot;,&quot;prompt&quot;:&quot;Minimal surreal isometric 3D editorial poster...&quot;}\n{&quot;out&quot;:&quot;02-flow.png&quot;,&quot;prompt&quot;:&quot;Minimal surreal isometric 3D editorial poster...&quot;}\n</code></pre>\n<p>生成原图：</p>\n<pre><code class=\"language-bash\">imgasset generate prompts.jsonl \\\n  --raw-dir temp/imgasset/article/raw \\\n  --skip-existing\n</code></pre>\n<p>这里的 <code>--skip-existing</code> 很实用。图片生成经常比较慢，也可能遇到网络波动。一批图如果生成到第三张失败，下一次重跑时不应该把前两张再生成一遍。</p>\n<h2>把两步连起来</h2>\n<p>单独有 <code>tinify-cli</code> 和 <code>imgasset generate</code> 已经能覆盖大部分场景，但最顺手的还是一条命令跑完整流程。</p>\n<pre><code class=\"language-bash\">imgasset run prompts.jsonl \\\n  --raw-dir temp/imgasset/article/raw \\\n  --publish-dir public/assets/article \\\n  --format jpeg \\\n  --background white \\\n  --skip-existing\n</code></pre>\n<p>这个命令会先生成原图，再调用内置依赖里的 <code>tinify-cli</code> 做压缩和格式转换。</p>\n<p><img src=\"/assets/posts/2026/image-asset-pipeline/03-generation-publish-flow.jpg\" alt=\"提示词、原图目录和发布目录被串成连续的图片生产路径\"></p>\n<p>也就是说，一个项目只要安装 <code>imgasset</code>，就同时拥有了生图和压缩能力，不需要再单独维护一套压缩脚本。</p>\n<p>实际使用时，我通常会让目录结构保持这样：</p>\n<pre><code class=\"language-text\">temp/\n  imgasset/\n    my-article/\n      raw/\npublic/\n  assets/\n    posts/\n      2026/\n        my-article/\nprompts.jsonl\n</code></pre>\n<p><code>temp/</code> 放原图和临时文件，进入 <code>.gitignore</code>。</p>\n<p><code>public/assets/</code> 放压缩后的发布图，可以被文章引用。</p>\n<p>Markdown 里只引用最终输出：</p>\n<pre><code class=\"language-markdown\">![内容系统示意图](/assets/posts/2026/my-article/01-context.jpg)\n</code></pre>\n<p>这个边界很重要。原图是生产资料，发布图才是网站资产。</p>\n<h2>为什么不用一个脚本解决</h2>\n<p>最开始当然可以用一个脚本解决。</p>\n<p>甚至对一个项目来说，一个脚本往往是最省事的。把 API key 从环境变量里读出来，循环请求图片接口，生成后再调用压缩 API，几十行代码就能跑起来。</p>\n<p>问题在于，这类脚本很容易变成一次性资产。</p>\n<p>当第二个项目也需要类似流程时，就会开始复制脚本。复制之后又会改目录、改模型、改压缩格式、改代理、改错误处理。再过一段时间，就很难判断哪一份才是最新实践。</p>\n<p>工具化的价值不在于代码量更少，而在于把边界固定下来：</p>\n<ul>\n<li>API key 永远不进项目。</li>\n<li>原图默认进入临时目录。</li>\n<li>提示词用 JSONL 保存。</li>\n<li>输出路径由命令或项目配置决定。</li>\n<li>压缩能力通过依赖提供，不要求每个项目额外安装。</li>\n<li>中断后可以继续跑。</li>\n</ul>\n<p>这些约定一旦稳定下来，后续每个项目都能复用同一套工作流。</p>\n<h2>关于安全和开源</h2>\n<p>这两个工具都发布到了 npm，也放到了 GitHub 上。</p>\n<p>开源前我重点处理了三件事。</p>\n<p>第一是密钥。API key 只能存在全局 secret 文件或环境变量里，项目配置、提示词文件、日志和报告都不应该包含密钥。</p>\n<p><img src=\"/assets/posts/2026/image-asset-pipeline/04-config-secret-boundary.jpg\" alt=\"配置、密钥和项目资产之间保持清晰边界\"></p>\n<p>第二是示例。示例里只能出现 <code>https://api.example.com/v1</code> 这种占位 base URL，不能把任何实际使用的服务地址写进去。</p>\n<p>第三是发布流程。两个包都尽量走标准 npm 包形态，<code>imgasset</code> 还配置了 GitHub Actions 和 npm Trusted Publishing。发版时只需要：</p>\n<pre><code class=\"language-bash\">pnpm run release\n</code></pre>\n<p>脚本会递增版本号、打 tag，GitHub Actions 再根据 tag 发布到 npm。</p>\n<p>这套流程看起来比手动 <code>npm publish</code> 麻烦一点，但长期更可靠。尤其是开源包，发布过程越可追溯越好。</p>\n<h2>最后还是回到写文章</h2>\n<p>这套工具做完之后，变化最明显的地方不是少敲了几行命令，而是配图不再打断写作节奏。</p>\n<p>以前一想到要给文章补图，脑子里会先冒出一串杂事：提示词放哪，原图放哪，压缩后叫什么名字，路径会不会写错，失败后要从哪里接着跑。每件事都很小，但它们会把注意力从文章里拉出来。</p>\n<p>现在流程更接近这样：</p>\n<ol>\n<li>读文章，决定需要几张图。</li>\n<li>写 <code>prompts.jsonl</code>。</li>\n<li>跑 <code>imgasset run</code>。</li>\n<li>把输出图插进 Markdown。</li>\n<li>构建检查。</li>\n</ol>\n<p>真正需要判断的部分还在：文章适合什么意象，几张图够不够，图片放在哪里能帮助阅读，哪张图虽然好看但不该用。工具只是把目录、命令、格式转换和失败重试这些事情固定下来。</p>\n<p>所以这篇文章想记录的不是“又写了两个 npm 包”，而是一个小流程变稳的过程。</p>\n<p>小工具最有价值的状态，大概就是平时感觉不到它的存在，但换一个项目时又能马上带走。<code>tinify-cli</code> 处理压缩这一步，<code>imgasset</code> 把生图、原图保存和发布输出串起来。它们加在一起，解决的不是某一次图片生成，而是下一篇文章、下一个站点、下一个项目里仍然能复用的图片资产流程。</p>\n","date_published":"2026-05-17T00:00:00.000Z","tags":["图片处理","AI","CLI","npm","工程化"],"language":"zh"},{"id":"https://www.lihuanyu.com/en/posts/2026/static-sites-renaissance-why-i-chose-astro/","url":"https://www.lihuanyu.com/en/posts/2026/static-sites-renaissance-why-i-chose-astro/","title":"Static Sites Are Having a Renaissance. Here Is Why I Chose Astro","summary":"A reflection on why Astro fits content-driven static sites, based on recent website projects, a blog migration from Hexo to Astro, and comparisons with Hugo, Eleventy, VitePress, Docusaurus, and Next.js.","content_html":"<p>Cloudflare recently <a href=\"https://x.com/Cloudflare/status/2050139665517199704\">asked a question on X</a>:</p>\n<blockquote>\n<p>Static sites are having a renaissance. What is your favorite static site generator right now and why do you prefer it?</p>\n</blockquote>\n<p>Astro appeared often in the replies.</p>\n<p>That matches my own experience over the past year. I have been using Astro in quite a few new projects, mostly official websites, product pages, and content sites. This blog also moved from Hexo to a new system built on top of Astro.</p>\n<p>I increasingly feel that static sites are not outdated. They are becoming important again.</p>\n<p>But this round of static sites is not a return to old template systems, plugin piles, and theme tweaking. It is a way to use a modern JavaScript toolchain during development while shipping standard HTML, CSS, and only the JavaScript that is actually needed.</p>\n<p>Astro fits exactly into that position.</p>\n<p><a href=\"/posts/2026/%E9%9D%99%E6%80%81%E7%BD%91%E7%AB%99%E5%9B%9E%E6%BD%AE%E6%97%B6-%E6%88%91%E4%B8%BA%E4%BB%80%E4%B9%88%E9%80%89%E6%8B%A9-Astro/\">Chinese version of this article</a></p>\n<p><img src=\"/assets/posts/2026/astro-static-sites/01-static-renaissance.jpg\" alt=\"Static content pages becoming central again in a modern toolchain\"></p>\n<h2>Why Static Sites Matter Again</h2>\n<p>The advantages of static sites have always been there: they are fast, stable, cheap to host, easy to cache, simple to deploy, and have a small security surface.</p>\n<p>If a page mainly exists to present content, such as a blog post, homepage, product page, documentation page, portfolio, or campaign page, the ideal output should usually be HTML and CSS. Users should not have to download a large JavaScript bundle first and then wait for the browser to assemble the actual content on the client side.</p>\n<p>Over the past few years, the frontend community has become used to building everything as if it were a Web app. React, Vue, Next.js, Nuxt, and SvelteKit are all powerful, but their mental model often starts from “application”: state, routing, data fetching, hydration, client-side interaction, server rendering, caching, and edge runtime.</p>\n<p>Those capabilities matter for complex applications. For many content-driven websites, however, they are not the starting point. They are extra cost.</p>\n<p>The renaissance of static sites is not nostalgia. It is a practical judgment: if a website is mainly content, the content should be delivered directly to the browser first.</p>\n<h2>Astro’s Core Appeal</h2>\n<p><a href=\"https://docs.astro.build/en/concepts/why-astro/\">Astro’s own documentation</a> describes it as a Web framework for content-driven websites. It is aimed at blogs, marketing sites, ecommerce content pages, documentation, portfolios, community sites, and similar use cases.</p>\n<p>That positioning matters. Astro is not trying to cover every Web shape first and then offer static export as a secondary feature. It starts from content-driven sites.</p>\n<p>For me, Astro’s appeal is this combination:</p>\n<ul>\n<li>During development, it feels like modern frontend engineering: components, TypeScript, Vite, Markdown, MDX, and the npm ecosystem.</li>\n<li>After build, the output is a high-performance static site made of standard HTML and CSS.</li>\n<li>JavaScript is not the foundation of every page by default. It is progressive enhancement for interaction.</li>\n</ul>\n<p>This is very different from traditional static site generators. Hexo, Jekyll, and Hugo can all turn Markdown into static HTML, but their development experience feels closer to a content system or a template system. Astro feels more like modern frontend engineering, without forcing the final output to carry the complexity of a frontend application.</p>\n<p>That balance is comfortable.</p>\n<h2>Less JavaScript by Default</h2>\n<p>Astro is most often associated with islands architecture.</p>\n<p>In <a href=\"https://docs.astro.build/en/concepts/islands/\">Astro’s explanation</a>, most of a page is rendered as static HTML. Only the areas that need interactivity or personalization run as JavaScript islands. By default, Astro components output HTML and CSS. They do not send a client-side runtime to the browser.</p>\n<p><img src=\"/assets/posts/2026/astro-static-sites/02-islands-less-javascript.jpg\" alt=\"Most of the page stays static while a few interactive areas become JavaScript islands\"></p>\n<p>The value here is not only better performance.</p>\n<p>It changes the default assumption of frontend development back to something more reasonable: a page should be a document first, and only become an application where necessary.</p>\n<p>In my own Astro usage, I actually do not use many React or Vue components inside Astro yet. I also do not use islands heavily. Many pages are just Markdown, Astro components, CSS, and a small amount of script. That does not weaken Astro’s value. It proves the default model is right.</p>\n<p>On many official websites, very little client-side JavaScript is truly necessary: navigation menus, theme toggles, form validation, carousels, search boxes, and a few animations. Turning the whole site into a client-side app is not always a good tradeoff.</p>\n<p>Astro’s strength is that one interactive component does not force the whole page to carry the cost of an application framework.</p>\n<h2>Content Is a First-Class Concern</h2>\n<p>Another important strength of Astro is its content model.</p>\n<p><a href=\"https://docs.astro.build/en/guides/content-collections/\">Content Collections</a> make Markdown, MDX, JSON, and other content sources work with schemas, type hints, and validation. For blogs, documentation, official websites, and product content pages, this is much more reliable than simply walking through a folder of files.</p>\n<p><img src=\"/assets/posts/2026/astro-static-sites/03-content-collections.jpg\" alt=\"Markdown, MDX, JSON, RSS, search indexes, and sitemaps organized as a content system\"></p>\n<p>When maintaining a content-driven site for the long term, rendering a page is usually not the hard part. The harder parts are questions like these:</p>\n<ul>\n<li>Are all article fields complete?</li>\n<li>Are tags, categories, dates, and summaries consistent?</li>\n<li>How should multilingual content be organized?</li>\n<li>How should RSS, sitemaps, and search indexes be generated?</li>\n<li>How should old URLs remain compatible?</li>\n<li>How should images, code blocks, and external links stay maintainable over time?</li>\n</ul>\n<p>Astro does not solve every content governance problem automatically, but it provides a better modern engineering base. Content is not merely attached to a template system. It becomes an input that can be handled by the type system, build process, and component model together.</p>\n<p>When this blog moved away from Hexo, this was one of the things I cared about most. Markdown remained the content source, but the build, routes, feeds, search, <code>llms.txt</code>, multilingual structure, and deployment around Markdown could be organized in a more modern way.</p>\n<h2>Static First, But Not Static Only</h2>\n<p>Astro is easy to understand as a static site generator, but it has already moved beyond the traditional meaning of SSG.</p>\n<p><a href=\"https://astro.build/blog/astro-6/\">Astro 6.0</a> was released in March 2026. It brought a built-in Fonts API, a stable Content Security Policy API, Live Content Collections, and a reworked dev server and build pipeline. Live Content Collections allow content from external CMSs or APIs to be fetched at request time, so not every content change has to trigger a full rebuild.</p>\n<p>This shows that Astro is not stopping at “compile Markdown into HTML.” It remains static-first, but it leaves a path toward dynamic content, server rendering, and edge runtimes.</p>\n<p>That direction matters for content-driven websites.</p>\n<p>Many websites start out static, then gradually grow dynamic needs: subscription forms, user state, A/B tests, personalized recommendations, CMS updates, protected content, or small pieces of backend data. Starting with a full application framework can be too expensive early on. Choosing a purely static generator can make later expansion awkward.</p>\n<p>Astro sits between the two: make the static pages good first, then add dynamic capabilities only where they are needed.</p>\n<h2>The Cloudflare Signal</h2>\n<p>Another important change is Astro’s relationship with Cloudflare.</p>\n<p>In January 2026, Cloudflare published <a href=\"https://blog.cloudflare.com/astro-joins-cloudflare/\">Astro is joining Cloudflare</a>. Astro Technology Company joined Cloudflare, while Astro remained open source, MIT licensed, publicly governed, and committed to platform-agnostic deployment.</p>\n<p>That is a positive signal for Astro’s long-term value.</p>\n<p>Content-driven websites naturally fit Cloudflare’s infrastructure. Static assets, CDN, edge cache, Workers, Pages, R2, and D1 all point in the same general direction: make websites faster, closer to users, and easier to deploy. If Astro continues to improve Cloudflare runtime support while staying platform-agnostic, that is a good position for developers.</p>\n<p><img src=\"/assets/posts/2026/astro-static-sites/04-cloudflare-edge-path.jpg\" alt=\"A content site moving closer to visitors through edge caching and deployment nodes\"></p>\n<p>Of course, joining a large company does not automatically make a framework better. For an open source project, governance, community, roadmap, and real usage experience still matter most. But Cloudflare choosing Astro at least shows that content-driven sites, static-first architecture, and edge deployment are not niche directions.</p>\n<h2>How It Compares with Other Options</h2>\n<p>Is Astro the first choice for static sites today?</p>\n<p>My answer is: if the site is content-driven and the developers mainly work in the JavaScript / TypeScript ecosystem, Astro can be the first choice. But it is not the only best answer for every static site.</p>\n<p><img src=\"/assets/posts/2026/astro-static-sites/05-choose-static-tool.jpg\" alt=\"Different static site tools branching from the same content decision square\"></p>\n<p><a href=\"https://gohugo.io/\">Hugo</a> is still very strong. It builds fast, ships as a single binary, and works well for large Markdown-heavy sites with little frontend customization. If a team does not need modern frontend components and mainly wants a stable, fast static generator, Hugo remains highly competitive.</p>\n<p><a href=\"https://www.11ty.dev/\">Eleventy</a> is more plain and closer to traditional templating. It has less framework feel and suits developers who want direct control over HTML output while keeping distance from frontend runtimes.</p>\n<p><a href=\"https://vitepress.dev/\">VitePress</a> and <a href=\"https://docusaurus.io/\">Docusaurus</a> are better fits for documentation sites. Versioning, sidebars, documentation navigation, search, code blocks, and theme conventions are important in documentation. For product docs, using a dedicated documentation framework is often more efficient.</p>\n<p><a href=\"https://nextjs.org/docs/app/building-your-application/deploying/static-exports\">Next.js</a>, Nuxt, and SvelteKit can also export static sites, but they are more natural for application-style websites. If a project has login state, dashboards, complex forms, real-time data, heavy interaction, or server-side data flows, an application framework is usually a better fit. Static export is one of their capabilities, but it is not their cleanest starting mental model.</p>\n<p>Traditional blog systems such as Hexo and Jekyll are still usable. They are mature, stable, rich in themes, and have clear migration paths. But for developers used to modern frontend engineering, Astro’s development experience, component model, and room for extension are easier to live with.</p>\n<p>So the choice should not be based on which framework is the loudest. A better rule is:</p>\n<ul>\n<li>If the site is mostly content, consider Astro, Hugo, or Eleventy first.</li>\n<li>If it is mostly documentation, consider VitePress, Docusaurus, or Starlight first.</li>\n<li>If it is mostly an application, consider Next.js, Nuxt, or SvelteKit first.</li>\n<li>If it is mostly a blog and you want both modern frontend development and static output, Astro deserves serious consideration.</li>\n</ul>\n<h2>What Astro Is Not For</h2>\n<p>Astro’s strength comes from being content-first. That also means it is not the best choice for every scenario.</p>\n<p>If a project is essentially an admin system, SaaS application, collaboration tool, complex editor, data dashboard, or anything that needs a large amount of client-side state from the homepage onward, Astro may not be the most natural choice. It can use React, Vue, and Svelte. It can also do server rendering. But the complexity of those projects usually does not live in “static page output.” It lives in application state, data flow, permissions, forms, real-time collaboration, and interaction structure.</p>\n<p>In those cases, Next.js, Nuxt, SvelteKit, or even a traditional backend full-stack framework may be a better fit.</p>\n<p>One common misunderstanding is to treat Astro as a “faster React framework.” I do not think that is accurate. Astro is closer to a content website framework. It allows React to be embedded where needed, but its core value is not making React faster. Its core value is making most pages not become React applications in the first place.</p>\n<p>That distinction matters.</p>\n<h2>Why I Chose Astro</h2>\n<p>Moving from Hexo to Astro did not feel like switching to a trendier tool. It made the engineering model of static websites clearer.</p>\n<p>With older static blog systems, I often felt a wall between the content system and frontend engineering. Writing posts, adjusting themes, editing templates, adding plugins, and handling builds often felt like working inside a relatively closed ecosystem.</p>\n<p>Astro feels different. It respects the nature of static websites, but brings the development experience back into modern frontend engineering:</p>\n<ul>\n<li>Pages are components.</li>\n<li>Styles can be organized locally.</li>\n<li>Markdown is the content source.</li>\n<li>Builds are powered by Vite.</li>\n<li>Client-side JavaScript is introduced only when interaction needs it.</li>\n<li>Server or edge runtime capabilities can be added when dynamic behavior needs them.</li>\n</ul>\n<p>That is why I keep choosing Astro for new official website projects.</p>\n<p>For official websites and blogs, the most important thing is not that the technology stack looks complete. It is that users can open pages quickly, search engines can understand the content, content stays maintainable over time, deployment remains simple, and migration cost stays manageable.</p>\n<p>Astro gives a direct answer to those needs.</p>\n<h2>Static Sites Are Not a Step Back</h2>\n<p>The return of static sites is not a regression in frontend engineering.</p>\n<p>It is closer to the frontend community rediscovering that the Web’s basic capabilities are already strong. HTML can express content. CSS can handle a large amount of presentation. Browsers can render pages directly. CDNs can distribute content globally.</p>\n<p>JavaScript is important, but it should not automatically become the foundation of every page.</p>\n<p>This is where Astro’s value sits. It does not reject modern frontend development, and it does not push everything into a client-side application. It simply gives a better default:</p>\n<p>Ship the page first. Enhance only where necessary.</p>\n<p>For content-driven static websites, that is almost the simplest and most effective engineering principle.</p>\n<p>So if I were building a blog, official website, portfolio, product landing page, marketing site, or content portal today, I would consider Astro first. Not because it is the most fashionable option, but because its default tradeoffs match the real needs of those sites.</p>\n<p>Static sites did not disappear. They finally got a toolchain that feels right for modern developers.</p>\n","date_published":"2026-05-16T00:00:00.000Z","tags":["Astro","Static Sites","Frontend","Engineering"],"language":"en"},{"id":"https://www.lihuanyu.com/posts/2026/%E9%9D%99%E6%80%81%E7%BD%91%E7%AB%99%E5%9B%9E%E6%BD%AE%E6%97%B6-%E6%88%91%E4%B8%BA%E4%BB%80%E4%B9%88%E9%80%89%E6%8B%A9-Astro/","url":"https://www.lihuanyu.com/posts/2026/%E9%9D%99%E6%80%81%E7%BD%91%E7%AB%99%E5%9B%9E%E6%BD%AE%E6%97%B6-%E6%88%91%E4%B8%BA%E4%BB%80%E4%B9%88%E9%80%89%E6%8B%A9-Astro/","title":"静态网站回潮时，我为什么选择 Astro","summary":"从 Cloudflare 关于静态站点生成器的讨论出发，结合个人官网项目和博客从 Hexo 迁移到 Astro 的实践，讨论 Astro 为什么适合内容型静态网站，以及它和 Hugo、Eleventy、VitePress、Docusaurus、Next.js 等方案的取舍。","content_html":"<p>最近 <a href=\"https://x.com/Cloudflare/status/2050139665517199704\">Cloudflare 在 X 上问了一个问题</a>：</p>\n<blockquote>\n<p>Static sites are having a renaissance. What is your favorite static site generator right now and why do you prefer it?</p>\n</blockquote>\n<p>评论区里 Astro 的名字出现得很频繁。</p>\n<p>这和我最近一段时间的感受很接近。过去一年，我在不少新项目里使用 Astro，主要是官网、产品介绍页、内容站这类偏静态的网站。这个博客也从 Hexo 迁移到了基于 Astro 的新系统。</p>\n<p>我越来越觉得，静态网站并没有过时。相反，它在重新变得重要。</p>\n<p>只不过这一次的静态网站，不再是早年那种模板、插件和主题拼出来的站点，而是用现代 JavaScript 工具链开发，最终交付标准 HTML、CSS 和少量必要 JavaScript 的网站。</p>\n<p>Astro 刚好踩中了这个位置。</p>\n<p><a href=\"/en/posts/2026/static-sites-renaissance-why-i-chose-astro/\">English version: Static Sites Are Having a Renaissance. Here Is Why I Chose Astro</a></p>\n<p><img src=\"/assets/posts/2026/astro-static-sites/01-static-renaissance.jpg\" alt=\"静态内容页面在现代工具链中重新成为网站中心\"></p>\n<h2>静态网站为什么又值得重视</h2>\n<p>静态网站的优势一直都在：快、稳定、便宜、容易缓存、部署简单、安全面小。</p>\n<p>如果一个页面的主要任务是展示内容，比如博客文章、官网首页、产品介绍、文档、作品集、活动页，那么最理想的交付物本来就应该是 HTML 和 CSS。用户点开页面时，不应该先下载一大包 JavaScript，再等浏览器在客户端把内容组装出来。</p>\n<p>过去几年，前端社区习惯了用构建 Web App 的方式构建一切。React、Vue、Next.js、Nuxt、SvelteKit 都很强，但这些框架的心智模型经常从“应用”开始：状态、路由、数据请求、hydration、客户端交互、服务端渲染、缓存、边缘运行时。</p>\n<p>这些能力对复杂应用很重要。但对很多内容型网站来说，它们不是起点，而是额外成本。</p>\n<p>静态网站回潮，本质上不是怀旧，而是一个更务实的判断：如果网站主要是内容，那就应该优先把内容直接交付给浏览器。</p>\n<h2>Astro 的核心吸引力</h2>\n<p><a href=\"https://docs.astro.build/en/concepts/why-astro/\">Astro 官方文档</a>对自己的定位很明确：它是面向内容驱动网站的 Web 框架，适合博客、营销站、电商内容页、文档、作品集、社区站等场景。</p>\n<p>这句话很重要。Astro 不是试图覆盖所有 Web 形态，然后顺便支持静态导出。它一开始就把内容型网站作为核心目标。</p>\n<p>对我来说，Astro 最吸引人的地方是这种组合：</p>\n<ul>\n<li>编写时是现代前端体验，可以用组件、TypeScript、Vite、Markdown、MDX、npm 生态。</li>\n<li>构建后是性能很好的静态产物，HTML、CSS 都是标准的。</li>\n<li>JavaScript 默认不是页面的基础设施，而是交互增强。</li>\n</ul>\n<p>这和传统静态站点生成器的差别很明显。Hexo、Jekyll、Hugo 这些工具都能把 Markdown 变成静态 HTML，但它们的开发体验更接近“内容系统”或“模板系统”。Astro 则更像现代前端工程，但产物又没有被前端应用的复杂度拖住。</p>\n<p>这个平衡点很舒服。</p>\n<h2>默认少发 JavaScript</h2>\n<p>Astro 最常被提到的是岛屿架构。</p>\n<p>在 <a href=\"https://docs.astro.build/en/concepts/islands/\">Astro 的解释</a>里，页面的大部分内容会渲染成静态 HTML，只有需要交互或个性化的区域才作为 JavaScript 岛屿运行。默认情况下，Astro 组件会输出 HTML 和 CSS，不会把客户端运行时一起发给浏览器。</p>\n<p><img src=\"/assets/posts/2026/astro-static-sites/02-islands-less-javascript.jpg\" alt=\"大部分页面保持静态，只有少量交互区域成为 JavaScript 岛屿\"></p>\n<p>这件事的价值不只是“性能更好”。</p>\n<p>它更像是把前端开发的默认值改了回来：页面应该先是文档，然后才在必要处变成应用。</p>\n<p>我现在使用 Astro 时，其实很少用 React、Vue 这些框架组件，也还没有大量使用岛屿能力。很多页面就是 Markdown、Astro 组件、CSS 和少量脚本。但这并不影响 Astro 的价值。恰恰相反，Astro 对纯静态页面很友好，说明它的默认模型足够正确。</p>\n<p>在很多官网项目里，真正需要客户端 JavaScript 的地方很少：导航菜单、主题切换、表单校验、图片轮播、搜索框、少量动画。把整站做成一个客户端应用，并不总是合理。</p>\n<p>Astro 的好处在于，它不会因为页面上有一个交互组件，就要求整页都背上应用框架的成本。</p>\n<h2>内容是第一等公民</h2>\n<p>Astro 另一个重要长处是内容能力。</p>\n<p><a href=\"https://docs.astro.build/en/guides/content-collections/\">Content Collections</a> 让 Markdown、MDX、JSON 等内容可以有 schema、类型提示和校验。对博客、文档、官网、产品内容页来说，这比单纯遍历文件目录可靠得多。</p>\n<p><img src=\"/assets/posts/2026/astro-static-sites/03-content-collections.jpg\" alt=\"Markdown、MDX、JSON、RSS、搜索索引和站点地图被组织成内容系统\"></p>\n<p>一个内容型网站长期维护时，真正麻烦的通常不是把页面渲染出来，而是这些事情：</p>\n<ul>\n<li>文章字段是否完整。</li>\n<li>标签、分类、日期、摘要是否规范。</li>\n<li>多语言内容如何组织。</li>\n<li>RSS、sitemap、搜索索引如何生成。</li>\n<li>旧链接如何兼容。</li>\n<li>图片、代码块、外链如何长期可维护。</li>\n</ul>\n<p>Astro 不能自动解决所有内容治理问题，但它给了一个更适合现代工程的基础。内容不是模板系统的附属品，而是可以被类型系统、构建流程和组件系统共同处理的输入。</p>\n<p>这个博客从 Hexo 迁移出来时，我最看重的也是这一点。Markdown 仍然是内容源，但围绕 Markdown 的构建、路由、feed、搜索、<code>llms.txt</code>、多语言和部署，可以用更现代的方式组织。</p>\n<h2>静态优先，但不是只能静态</h2>\n<p>Astro 很容易被理解成“静态站点生成器”，但它已经不只是传统意义上的 SSG。</p>\n<p><a href=\"https://astro.build/blog/astro-6/\">Astro 6.0</a> 发布于 2026 年 3 月，带来了内置 Fonts API、稳定的 Content Security Policy API、Live Content Collections，并且重构了 dev server 和构建流水线。Live Content Collections 让外部 CMS 或 API 内容可以在请求时获取，不必每次内容变化都重新构建。</p>\n<p>这说明 Astro 的演进方向并不是停留在“把 Markdown 编译成 HTML”。它仍然静态优先，但保留了向动态内容、服务端渲染、边缘运行时扩展的路径。</p>\n<p>这个方向对内容型网站很关键。</p>\n<p>因为很多网站一开始都是静态的，后来会慢慢长出一些动态需求：订阅表单、用户状态、A/B 测试、个性化推荐、CMS 实时更新、受保护内容、局部后台数据。直接从全栈应用框架开始，早期成本偏高；只选一个纯静态生成器，后续扩展又可能受限。</p>\n<p>Astro 的位置介于两者之间：先把静态页面做好，再让动态能力按需出现。</p>\n<h2>Cloudflare 的信号</h2>\n<p>还有一个值得关注的变化是 Astro 和 Cloudflare 的关系。</p>\n<p>2026 年 1 月，Cloudflare 发布了 <a href=\"https://blog.cloudflare.com/astro-joins-cloudflare/\">Astro is joining Cloudflare</a>：Astro Technology Company 加入 Cloudflare，Astro 继续保持开源、MIT 许可、公开路线图和开放治理，也继续强调跨平台部署。</p>\n<p>这件事对 Astro 的长期价值是加分项。</p>\n<p>内容型网站和 Cloudflare 的基础设施天然匹配。静态资源、CDN、边缘缓存、Workers、Pages、R2、D1，这些能力都围绕一个方向展开：让网站更快、更靠近用户、更容易部署。Astro 如果继续强化 Cloudflare 运行时支持，同时保持平台无关，对开发者来说是一个比较好的状态。</p>\n<p><img src=\"/assets/posts/2026/astro-static-sites/04-cloudflare-edge-path.jpg\" alt=\"内容站点通过边缘缓存和部署节点靠近不同地区的访问者\"></p>\n<p>当然，框架加入大公司并不自动等于更好。开源项目最重要的仍然是治理、社区、路线图和实际使用体验。但 Cloudflare 选择 Astro，至少说明内容型网站、静态优先和边缘部署这条路线并不边缘。</p>\n<h2>和其他方案怎么选</h2>\n<p>Astro 是不是现在做静态网站的第一选择？</p>\n<p>我的答案是：如果是内容型网站，并且开发者主要在 JavaScript / TypeScript 生态里工作，Astro 可以作为第一选择。但它不是所有静态网站的唯一最优解。</p>\n<p><img src=\"/assets/posts/2026/astro-static-sites/05-choose-static-tool.jpg\" alt=\"不同静态站点工具从同一个内容决策广场分出不同路径\"></p>\n<p><a href=\"https://gohugo.io/\">Hugo</a> 仍然非常强。它构建速度快，单文件二进制，适合大量 Markdown 内容和很少前端定制的站点。如果团队不需要现代前端组件，只想稳定、快速、长期生成静态页面，Hugo 很有竞争力。</p>\n<p><a href=\"https://www.11ty.dev/\">Eleventy</a> 更朴素，也更接近传统模板系统。它的框架感更弱，适合喜欢直接控制 HTML 输出、对前端运行时保持距离的开发者。</p>\n<p><a href=\"https://vitepress.dev/\">VitePress</a> 和 <a href=\"https://docusaurus.io/\">Docusaurus</a> 更适合文档站。版本管理、侧边栏、文档导航、搜索、代码块、主题约定，这些能力在文档场景里很重要。做产品文档时，选专门的文档框架通常更省心。</p>\n<p><a href=\"https://nextjs.org/docs/app/building-your-application/deploying/static-exports\">Next.js</a>、Nuxt、SvelteKit 这类框架也能做静态导出，但它们更适合应用型网站。如果项目有登录态、后台、复杂表单、实时数据、强交互、服务端数据流，直接使用应用框架会更自然。静态导出是它们的能力之一，但不是最清晰的心智起点。</p>\n<p>Hexo、Jekyll 这类传统博客系统也不是不能用。它们的优势是成熟、稳定、主题多、迁移路径清楚。只是对习惯现代前端工程的人来说，Astro 的开发体验、组件能力和扩展空间会更顺手。</p>\n<p>所以选择标准不应该是“哪个框架最火”，而应该是：</p>\n<ul>\n<li>如果主要是内容，优先 Astro、Hugo、Eleventy。</li>\n<li>如果主要是文档，优先 VitePress、Docusaurus、Starlight。</li>\n<li>如果主要是应用，优先 Next.js、Nuxt、SvelteKit。</li>\n<li>如果主要是博客，并且希望现代前端体验和静态产物兼得，Astro 很值得优先考虑。</li>\n</ul>\n<h2>Astro 不适合什么</h2>\n<p>Astro 的优势来自内容优先，也意味着它不是所有场景的最佳选择。</p>\n<p>如果一个项目本质上是后台系统、SaaS 应用、协同工具、复杂编辑器、数据看板，或者从首页开始就需要大量客户端状态，那么 Astro 未必是最自然的选择。它能接入 React、Vue、Svelte，也能做服务端渲染，但这类项目的复杂度通常不在“页面静态输出”，而在应用状态、数据流、权限、表单、实时协作和交互结构。</p>\n<p>这时 Next.js、Nuxt、SvelteKit，甚至传统后端全栈框架，可能更适合。</p>\n<p>Astro 的一个常见误区，是把它当成“更快的 React 框架”。我觉得这种理解并不准确。Astro 更像是一个内容网站框架。它允许在需要时嵌入 React，但它的核心价值不是让 React 更快，而是让多数页面不必先变成 React 应用。</p>\n<p>这个区别很重要。</p>\n<h2>我为什么选择 Astro</h2>\n<p>从 Hexo 到 Astro，我最大的感受不是“换了一个更潮的工具”，而是静态网站的工程模型变清楚了。</p>\n<p>以前做静态博客，常见感受是内容系统和前端工程之间有一道墙。写文章、调主题、改模板、加插件、处理构建，经常像是在一个偏封闭的生态里工作。</p>\n<p>Astro 的感觉不同。它仍然尊重静态网站的本质，但把开发体验带回现代前端：</p>\n<ul>\n<li>页面是组件。</li>\n<li>样式可以局部组织。</li>\n<li>Markdown 是内容源。</li>\n<li>构建由 Vite 驱动。</li>\n<li>需要交互时再引入客户端 JavaScript。</li>\n<li>需要动态能力时再接入服务端或边缘运行时。</li>\n</ul>\n<p>这也是我在新官网项目里反复选择 Astro 的原因。</p>\n<p>官网和博客最重要的不是“技术栈看起来完整”，而是用户打开得快、搜索引擎能读懂、内容长期可维护、部署链路简单、迁移成本可控。</p>\n<p>Astro 在这些点上给出的答案很直接。</p>\n<h2>静态网站不是退回去</h2>\n<p>静态网站的回潮，不是前端工程倒退。</p>\n<p>恰恰相反，它更像是前端社区在经历了多年应用框架膨胀之后，重新认识到 Web 的基础能力本来就很强。HTML 可以表达内容，CSS 可以完成大量视觉，浏览器可以直接渲染页面，CDN 可以把内容分发到全球。</p>\n<p>JavaScript 当然重要，但它不应该自动成为每个页面的地基。</p>\n<p>Astro 的价值就在这里：它没有否定现代前端，也没有把所有东西都推向客户端应用。它只是给了一个更合理的默认值：</p>\n<p>先输出网页，再按需增强。</p>\n<p>对内容型静态网站来说，这几乎就是最朴素、也最有效的工程原则。</p>\n<p>所以如果今天要做一个博客、官网、作品集、产品介绍页、营销站或者内容门户，我会优先考虑 Astro。不是因为它最流行，而是因为它的默认取舍和这类网站的真实需求一致。</p>\n<p>静态网站并没有消失。它只是终于等到了更适合现代开发者的工具。</p>\n","date_published":"2026-05-16T00:00:00.000Z","tags":["Astro","静态网站","前端","工程化"],"language":"zh"},{"id":"https://www.lihuanyu.com/posts/2026/AI%E6%97%B6%E4%BB%A3%E6%9C%89%E9%87%8D%E6%9E%84%E7%9A%84%E8%87%AA%E7%94%B1/","url":"https://www.lihuanyu.com/posts/2026/AI%E6%97%B6%E4%BB%A3%E6%9C%89%E9%87%8D%E6%9E%84%E7%9A%84%E8%87%AA%E7%94%B1/","title":"AI 时代，有重构的自由","summary":"从 Bun 迁移到 Rust 和个人项目从 SolidJS 迁到 Vue 出发，讨论 AI 如何降低重构成本，以及技术选型在 AI 时代为什么逐渐从一次性押注变成可持续修正。","content_html":"<p>过去做项目，最怕第一铲土挖错地方。</p>\n<p>语言、框架、目录结构、状态管理、部署方式，看上去是几项技术选择，实际常常是在给未来修路。路修对了，车跑得顺。路修歪了，车也能跑，只是每天多绕十公里，日子久了，司机会把绕路当成生活的一部分。</p>\n<p>软件项目有一种很顽固的惯性。</p>\n<p>今天的选择，会变成明天的依赖；明天的依赖，会变成后天的约束；约束再往后走，名字就改成了技术债。债这东西最厉害的地方，往往还不在代码里，而在人心里。大家都知道它别扭，也都知道最好改掉，可一想到要动，手又缩回去了。</p>\n<p>因为过去的重构太重。</p>\n<p>它很少像换一把椅子，更像给一栋已经住满人的楼换地基。窗户要留着，水电要通着，住户还不能被惊醒。很多团队最后选择的办法，是在墙上多钉几块木板。看起来加固了，实际只是让下一次维修更难下手。</p>\n<p>AI 时代，这种沉重感开始松动。</p>\n<p>程序员开始重新拿到一种久违的东西：重构的自由。</p>\n<p><a href=\"/en/posts/2026/developers-have-the-freedom-to-refactor-ai-era/\">English version: In the AI Era, Developers Have the Freedom to Refactor</a></p>\n<p><img src=\"/assets/posts/2026/refactor-freedom/01-technical-debt-map.jpg\" alt=\"技术债让软件项目的道路越来越难改\"></p>\n<h2>Bun 的一声响</h2>\n<p>2026 年 5 月 14 日，Bun 的 <a href=\"https://github.com/oven-sh/bun/pull/30412\">Rewrite Bun in Rust</a> PR 合并到了 main 分支。</p>\n<p>Bun 很长时间里给人的印象，是一个用 Zig 写出来的 JavaScript runtime。它快，锋利，有一点年轻工具特有的锐气。突然看到这样一个 PR，很难不愣一下：一个已经跑在大量开发者机器上的基础设施项目，居然把底层语言往 Rust 迁。</p>\n<p>这个 PR 不小。</p>\n<p>一百多万行新增，两千多个文件，六千多个提交。按老经验看，这种事像远征。路途长，粮草重，中间还容易掉队。换成很多商业项目，光是立项评审就够写几轮 PPT。</p>\n<p>把它写成“Zig 输了，Rust 赢了”，未免太省事。PR 说明里讲得清楚：代码库大体沿用原来的架构和数据结构，后续还会继续优化和清理，非 canary 版本要看官方发布节奏。</p>\n<p>更值得看的，是一个高速奔跑的工具，居然还有余力在底层材料上动刀。</p>\n<p>它不像推倒重建，更像给桥换钢材。桥的走向还在，受力图还在，通行目标也还在，只是过去容易生锈、容易断裂、维修费太高的地方，换成了另一种材料。</p>\n<p><img src=\"/assets/posts/2026/refactor-freedom/02-changing-foundation.jpg\" alt=\"在保持通行的同时更换底层材料\"></p>\n<p>过去这种事当然也能做。只是能做和做得起，中间隔着一条河。AI 把河面冻住了一部分，人终于可以试着走过去。</p>\n<p>这是一声很响的提醒：软件不必永远忍受自己的出生缺陷。</p>\n<h2>我的小后台</h2>\n<p>我自己最近也有一个小得多的例子。</p>\n<p>有个后台管理网页，最早用 SolidJS 写。SolidJS 的响应式模型很漂亮，写 demo 的时候也顺手。但真实业务不会只拿理念吃饭。后台系统要表格、表单、弹窗、筛选、权限、菜单、校验、导入导出，还要有足够多的组件和足够好找的答案。</p>\n<p>写着写着就发现，能做，慢。</p>\n<p>后台管理系统很少需要前端哲学。它要的是稳、快、省心。用户不会因为一个表单背后有精妙的响应式模型就多点一次保存。开发者也不会因为框架观念漂亮，就少写一个日期范围选择器。</p>\n<p>这类项目放在以前，大概率先忍着。</p>\n<p>因为迁移听起来麻烦。组件要搬，路由要搬，状态要搬，接口调用要搬，样式和细节也要搬。心里知道 Vue 生态更合适，手上还是会继续补丁。补着补着，项目也就老了。</p>\n<p>现在做法直接很多。</p>\n<p>把页面行为、接口形态、组件结构和关键业务逻辑整理清楚，让 AI 带着这些上下文往 Vue 迁。过程里当然还要检查，还要改，还要盯细节。但最沉的那部分体力活，已经有人帮着扛了。</p>\n<p><img src=\"/assets/posts/2026/refactor-freedom/03-migration-workbench.jpg\" alt=\"把旧页面拆成可以迁移的模块\"></p>\n<p>人需要花精力的地方，变成了判断。</p>\n<p>哪些行为必须一致，哪些旧写法可以丢掉，哪些地方应该趁迁移顺手整理，哪些地方最好原样保留。过去重构像搬砖，现在更像监工。砖还是砖，墙还是墙，但人的手终于不用一直陷在水泥里。</p>\n<p>技术选型也因此少了一点宿命感。</p>\n<p>以前选框架像早婚。合适不合适，都先过下去。现在更像阶段性合作。合适就继续，不合适就把账算清，把东西收拾好，然后换一条路。</p>\n<h2>反悔的手续费降下来了</h2>\n<p>AI 没有废掉架构。</p>\n<p>它废掉的是一部分对架构的迷信。</p>\n<p>过去许多选择之所以显得神圣，并非它们多么高明，只因改起来太累。改 import，改调用方式，改类型定义，改组件写法，补适配层，修一批又一批细碎错误。方向并不难看清，难的是走过去要踩一脚泥。</p>\n<p>AI 正好擅长这片泥地。</p>\n<p>相似模式的迁移、重复结构的改写、失败测试后的修补、跨文件的机械调整，这些事情以前会消耗大量心力。现在它们还会消耗时间，却不再那么可怕。</p>\n<p>人的注意力可以往上提一点。</p>\n<p>为什么迁？迁到哪里？成功的标准是什么？旧系统里哪些是业务规则，哪些只是历史包袱？哪些复杂度应该保留，哪些复杂度只是多年风沙堆出来的土坡？</p>\n<p>选择仍然有代价。</p>\n<p>AI 降低的是反悔的手续费。</p>\n<p>这点很重要。手续费下降以后，人可以更大胆地试错。新框架可以试，冷门方案也可以试，小项目可以先用最快的办法跑起来。早期技术选型不必像刻墓志铭一样慎重。</p>\n<p>可也别走到另一头。</p>\n<p>今天换框架，明天换语言，后天换数据库，把每一次新鲜感都包装成架构演进，那叫折腾。折腾久了，项目会像一间不断装修的房子，墙纸永远是新的，人却始终住不进去。</p>\n<h2>自由要打桩</h2>\n<p>重构自由有门槛。</p>\n<p>第一根桩是设计文档。</p>\n<p>文档要记下当时为什么这么做。代码能告诉人现在怎么跑，很难告诉人当初为什么绕了一个弯。很多看起来奇怪的实现，背后可能有业务限制、历史兼容、线上事故和一段没人想再提的夜班。</p>\n<p>第二根桩是测试。</p>\n<p>测试管行为。没有测试的大规模重构，就像夜里搬家，东西看着都装上车了，天亮才发现户口本和钥匙不见了。代码变漂亮，用户路径断掉，这种账最难算。</p>\n<p>第三根桩是业务分层。</p>\n<p>底层语言可以换，中间框架可以换，展示层可以换。业务规则最好别撒得到处都是。业务越集中，迁移越像搬家；业务越散，迁移越像考古。考古当然也能考，只是每挖一铲都怕碰碎东西。</p>\n<p>第四根桩是可回滚。</p>\n<p><img src=\"/assets/posts/2026/refactor-freedom/04-refactor-piles.jpg\" alt=\"文档、测试、边界和回滚托住重构自由\"></p>\n<p>重构在分支里跑通，只算一半。上线后不伤人，才算另一半。分阶段迁移、灰度、对照验证、日志观察、保留旧路径，这些办法看起来笨，却能保命。AI 能把施工队叫来，验收制度还得人自己建。</p>\n<p>有了这些桩，程序才有余地。</p>\n<p>文档在，测试在，边界在，版本记录在，第一次技术选型就不再像一道圣旨。它只是一个阶段性的决定。决定可以被尊重，也可以在证据充分时被修改。</p>\n<h2>架构从石碑变成草图</h2>\n<p>过去的架构像石碑。</p>\n<p>既然要刻下去，就希望它一开始足够正确，能挡风，能挨打，能撑很多年。石碑一旦刻错，改字很麻烦；整块推倒，又显得败家。</p>\n<p>AI 时代的架构更像草图。</p>\n<p><img src=\"/assets/posts/2026/refactor-freedom/05-architecture-sketch.jpg\" alt=\"架构从石碑变成可以修改的草图\"></p>\n<p>草图也要认真画。比例要对，边界要清，重点要明。可草图承认未来会变，承认今天掌握的信息不够，承认软件系统是一种活的生产工具。</p>\n<p>这会改变我们看新技术的眼神。</p>\n<p>过去遇到一个新框架，问题总是那些：成熟吗，生态大吗，维护者靠谱吗，三年后还在吗。这些问题仍然要问。只是还可以再加一个更现实的问题：</p>\n<p>以后迁得走吗？</p>\n<p>这个问题一出来，尺度就变了。</p>\n<p>冷门技术未必危险，热门技术也未必稳妥。好方案要看退出成本。它能保护业务逻辑，保护数据结构，保护外部契约，将来换掉也不至于伤筋动骨。坏方案哪怕今天很流行，只要把业务、框架、存储、构建和部署搅成一锅粥，也是在提前抵押未来。</p>\n<p>很多技术债，就是在一片掌声里借下来的。</p>\n<h2>规划的题目变了</h2>\n<p>初始规划仍然重要。</p>\n<p>只是题目换了。</p>\n<p>过去的问题是：第一次能不能选对？</p>\n<p>现在还要补一句：第一次没有完全选对，未来能不能改？</p>\n<p>这更接近真实世界。很多项目刚出生时，没人知道它以后会长成什么样。需求会变，团队会变，流量会变，商业模式会变，依赖的生态也会变。要求一个项目在第一天预见几年后的命运，有点像要求婴儿自己填写退休计划。</p>\n<p>更可靠的办法，是给未来留通道。</p>\n<p>接口边界清楚一点，领域模型干净一点，测试贴近真实行为一点，文档解释关键取舍一点，数据迁移方案保守一点。做这些事情并不显得时髦，却能让未来那个需要重构的人少骂几句。</p>\n<p>那个未来的人，多半还是自己。</p>\n<h2>技术债像账本</h2>\n<p>技术债不会消失。</p>\n<p>AI 消灭不了偷懒，消灭不了复杂业务，也消灭不了错误判断。只要软件还在现实世界里跑，债就会继续出现。区别在于，过去很多债像判决书，一盖章，人就被压住了；现在它更像账本，数额清楚，利息清楚，还款路径清楚，就有周转的可能。</p>\n<p>这对独立开发者尤其要紧。</p>\n<p>一个人做项目，最怕被早期选择困住。框架不合适，生态不顺手，架构越来越别扭，新功能写不动，旧代码不敢改。项目还没死，开发者先被自己的代码磨没了兴致。</p>\n<p>AI 给小团队和个人开发者多发了一张返程票。</p>\n<p>可以先用熟悉的方案把东西做出来。可以试一个新框架验证想法。可以在产品还小的时候换底座。也可以在业务长出新形态后，重新整理分层，少往旧结构上贴膏药。</p>\n<p>不过每一次心血来潮都喊重构，系统很快会被喊散。</p>\n<p>重构要让系统更接近业务本身，要降低未来变化的阻力，要把散乱的概念重新摆正。追新名词、换新皮肤、给简历添技术栈，这些事情可以做，别借重构的名义。</p>\n<h2>有自由，也要有纪律</h2>\n<p>AI 把一部分沉重的重复劳动从程序员肩上搬走了。</p>\n<p>这很要紧。</p>\n<p>程序员最宝贵的能力，从来都不是把同一类代码改一千遍。更重要的是判断什么值得改，什么该留下，什么只是暂时能用，什么以后会勒住脖子。</p>\n<p>有了 AI，重构少了一点悲壮感。它可以更日常，更频繁，也更像工程本来的样子：观察系统，发现问题，调整结构，验证行为，继续前进。</p>\n<p>自由需要纪律托住。</p>\n<p>敢试，也要会收拾；敢开工，也要敢推倒；不迷信第一次选择，也不轻慢长期结构。设计文档、测试用例、业务边界和发布纪律，就是这种自由的地基。</p>\n<p>以前的软件项目，常常像被第一次技术选型押上轨道的列车。轨道歪了，车也只能一路冒烟往前开。</p>\n<p>AI 时代，轨道终于没那么神圣了。</p>\n<p>路可以改，桥可以重修，车也可以换。目的地要清楚，沿途要有标记，每一次改道要经得起验证。做得到这些，程序员就不必把早年的选择当成一生的枷锁。</p>\n<p>AI 没有替程序员免去判断。</p>\n<p>它只是把那堵由重复劳动砌成的墙打矮了一点。墙矮了，人可以看见远处的路。</p>\n<p>看见了，还得自己走。</p>\n","date_published":"2026-05-14T00:00:00.000Z","tags":["AI","重构","架构","开发者"],"language":"zh"},{"id":"https://www.lihuanyu.com/en/posts/2026/developers-have-the-freedom-to-refactor-ai-era/","url":"https://www.lihuanyu.com/en/posts/2026/developers-have-the-freedom-to-refactor-ai-era/","title":"In the AI Era, Developers Have the Freedom to Refactor","summary":"A reflection on Bun's move toward Rust, a small SolidJS-to-Vue migration, and why technical choices in the AI era are becoming less like permanent bets and more like decisions that can be revised.","content_html":"<p>In the past, the first shovel of dirt on a software project felt unusually heavy.</p>\n<p>Language, framework, folder structure, state management, deployment model: they looked like technical choices. In practice, they often became roads for the future. Build the road well, and the car runs smoothly. Build it crooked, and the car still moves, but it takes a ten-kilometer detour every day. After long enough, the driver starts treating the detour as part of life.</p>\n<p>Software projects have stubborn inertia.</p>\n<p>Today’s choice becomes tomorrow’s dependency. Tomorrow’s dependency becomes the next day’s constraint. Give that constraint enough time, and it gets a more familiar name: technical debt. The sharpest part of debt is often not in the code. It is in people’s minds. Everyone knows the system is awkward. Everyone knows it would be better to fix it. But once someone thinks about actually touching it, the hand pulls back.</p>\n<p>Refactoring used to be heavy.</p>\n<p>It rarely felt like replacing a chair. It felt more like changing the foundation of a building that was already full of residents. The windows had to stay. The water and electricity had to keep running. Nobody inside was supposed to wake up. Many teams eventually chose the same solution: nail a few more boards onto the wall. It looked reinforced. It also made the next repair harder.</p>\n<p>In the AI era, that heaviness is starting to loosen.</p>\n<p>Developers are beginning to regain something that had been missing for a long time: the freedom to refactor.</p>\n<p><a href=\"/posts/2026/AI%E6%97%B6%E4%BB%A3%E6%9C%89%E9%87%8D%E6%9E%84%E7%9A%84%E8%87%AA%E7%94%B1/\">Chinese version of this article</a></p>\n<p><img src=\"/assets/posts/2026/refactor-freedom/01-technical-debt-map.jpg\" alt=\"Technical debt makes software roads harder to reroute\"></p>\n<h2>Bun Made a Loud Sound</h2>\n<p>On May 14, 2026, Bun’s <a href=\"https://github.com/oven-sh/bun/pull/30412\">Rewrite Bun in Rust</a> PR was merged into the main branch.</p>\n<p>For a long time, Bun was known as a JavaScript runtime written in Zig. It was fast, sharp, and had the edge of a young tool. Seeing a PR like that is hard to ignore: a piece of infrastructure already running on many developers’ machines was moving its lower-level language toward Rust.</p>\n<p>It was not a small PR.</p>\n<p>More than one million lines added. More than two thousand files touched. More than six thousand commits. By older engineering instincts, this looks like an expedition. The road is long, the supplies are heavy, and people may fall behind halfway. In many commercial projects, the approval process alone would generate several rounds of slides.</p>\n<p>It would be too easy to turn this into “Zig lost, Rust won.” The PR description is more careful. The codebase largely kept the same architecture and data structures. More optimization and cleanup work would follow. The non-canary release schedule still depends on official releases.</p>\n<p>The more interesting point is that a fast-moving tool still had room to cut into its own foundation.</p>\n<p>It did not look like burning the whole thing down. It looked more like replacing the steel in a bridge. The direction of the bridge remained. The load map remained. The goal of letting traffic pass remained. But the places that were easier to rust, crack, or cost too much to maintain could be rebuilt with a different material.</p>\n<p><img src=\"/assets/posts/2026/refactor-freedom/02-changing-foundation.jpg\" alt=\"Replacing the lower-level material while traffic keeps moving\"></p>\n<p>This kind of work was possible before. Possible and affordable are separated by a river. AI freezes part of that river, and people can finally try walking across.</p>\n<p>It is a loud reminder: software does not have to endure its birth defects forever.</p>\n<h2>My Small Admin UI</h2>\n<p>I recently had a much smaller version of the same feeling.</p>\n<p>I had an admin UI that was originally written in SolidJS. SolidJS has a beautiful reactive model, and it feels nice when writing demos. But real business does not eat ideas alone. An admin system needs tables, forms, dialogs, filters, permissions, menus, validation, imports, exports, enough components, and answers that are easy to find.</p>\n<p>After writing it for a while, the conclusion was simple: it could be done, but it was slow.</p>\n<p>Admin systems rarely need frontend philosophy. They need to be stable, fast, and low-friction. Users do not click Save one more time because a form is powered by an elegant reactive model. Developers do not write one less date-range picker because a framework has beautiful ideas.</p>\n<p>In the past, I would probably have endured it.</p>\n<p>Migration sounds troublesome. Components have to move. Routes have to move. State has to move. API calls have to move. Styles and small details have to move too. Even if Vue’s ecosystem clearly fits the admin UI better, it is easy to keep patching. Patch long enough, and the project gets old.</p>\n<p>Now the path is more direct.</p>\n<p>I organized the page behavior, API shape, component structure, and important business logic, then let AI migrate the project toward Vue with that context. The process still needed review, changes, and attention to detail. But the heaviest physical labor had someone else carrying it.</p>\n<p><img src=\"/assets/posts/2026/refactor-freedom/03-migration-workbench.jpg\" alt=\"Turning an old admin UI into movable migration pieces\"></p>\n<p>The human work shifted toward judgment.</p>\n<p>Which behaviors must remain identical? Which old patterns can be discarded? Which parts should be cleaned up during the migration? Which parts should be left unchanged? Refactoring used to feel like carrying bricks. Now it feels more like supervising the site. Bricks are still bricks. Walls are still walls. But human hands no longer have to stay buried in cement the whole time.</p>\n<p>Technical choices now feel a little less fated.</p>\n<p>Choosing a framework used to feel like an early marriage. Whether it was suitable or not, you kept living with it. Now it feels more like a temporary partnership. If it works, continue. If it does not, settle the accounts, pack the belongings, and take another road.</p>\n<h2>The Fee for Changing Your Mind Has Dropped</h2>\n<p>AI has not abolished architecture.</p>\n<p>It has abolished part of the superstition around architecture.</p>\n<p>Many old choices looked sacred not because they were brilliant, but because changing them was exhausting. Change imports. Change call sites. Change type definitions. Change component patterns. Add adapters. Fix one batch of small errors after another. The direction was often visible. The problem was the mud between here and there.</p>\n<p>AI is good at that mud.</p>\n<p>Migrating similar patterns, rewriting repetitive structures, fixing code after failed tests, and making mechanical cross-file adjustments used to consume a lot of energy. They still take time, but they no longer feel as frightening.</p>\n<p>Human attention can move a little higher.</p>\n<p>Why migrate? Migrate to what? What counts as success? Which parts of the old system are business rules, and which are historical baggage? Which complexity should stay, and which complexity is only a hill of dirt left by years of wind?</p>\n<p>Choices still have cost.</p>\n<p>AI lowers the fee for changing your mind.</p>\n<p>That matters. When the fee drops, people can try things more boldly. A new framework can be tested. A niche solution can be tested. A small project can start with the fastest path first. Early technical choices no longer have to be treated like inscriptions on a tombstone.</p>\n<p>But do not run to the other extreme.</p>\n<p>Change the framework today, the language tomorrow, and the database the day after, then call every appetite for novelty “architecture evolution.” That is just thrashing. Thrash long enough, and a project becomes a house under endless renovation. The wallpaper is always new. Nobody ever gets to live inside.</p>\n<h2>Freedom Needs Pilings</h2>\n<p>The freedom to refactor has requirements.</p>\n<p>The first piling is design documentation.</p>\n<p>Documentation should record why a choice was made. Code can tell people how the system runs now. It has a harder time explaining why someone took a strange turn years ago. Many odd implementations have business limits, historical compatibility, production incidents, or a night shift nobody wants to talk about behind them.</p>\n<p>The second piling is tests.</p>\n<p>Tests guard behavior. Large-scale refactoring without tests is like moving house at night. Everything seems to be loaded onto the truck. At dawn, the household register and the keys are missing. Beautiful code with broken user paths is the kind of bill nobody wants to pay.</p>\n<p>The third piling is business layering.</p>\n<p>The lower-level language can change. The middle framework can change. The presentation layer can change. Business rules should not be scattered everywhere. The more concentrated the business logic is, the more migration feels like moving house. The more scattered it is, the more migration feels like archaeology. Archaeology is possible, but every shovel may break something.</p>\n<p>The fourth piling is rollback.</p>\n<p><img src=\"/assets/posts/2026/refactor-freedom/04-refactor-piles.jpg\" alt=\"Documentation, tests, boundaries, and rollback support refactoring freedom\"></p>\n<p>Refactoring that works on a branch is only half done. The other half is going online without hurting people. Phased migration, canary release, comparison checks, log observation, and old path retention may look clumsy, but they keep systems alive. AI can bring in the construction crew. The acceptance system still has to be built by humans.</p>\n<p>With these pilings, a program has room to move.</p>\n<p>When documentation, tests, boundaries, and version history exist, the first technical choice no longer looks like an imperial decree. It is a decision for a stage. It can be respected. It can also be revised when there is enough evidence.</p>\n<h2>Architecture Becomes a Sketch</h2>\n<p>Architecture used to feel like a stone tablet.</p>\n<p>Once something was carved into it, people hoped it was correct enough, durable enough, and able to withstand years of weather. If the carving was wrong, changing the words was troublesome. Pushing the whole tablet over looked wasteful.</p>\n<p>Architecture in the AI era feels more like a sketch.</p>\n<p><img src=\"/assets/posts/2026/refactor-freedom/05-architecture-sketch.jpg\" alt=\"Architecture becomes a sketch that can be revised\"></p>\n<p>A sketch still deserves care. The proportions should be right. The boundaries should be clear. The emphasis should be visible. But a sketch admits that the future will change. It admits that today’s information is incomplete. It admits that a software system is a living production tool.</p>\n<p>This changes how we look at new technologies.</p>\n<p>In the past, the questions around a new framework were always familiar: Is it mature? Is the ecosystem large enough? Are the maintainers reliable? Will it still exist in three years? Those questions still matter. But one more practical question can be added:</p>\n<p>Can we leave later?</p>\n<p>Once that question appears, the scale changes.</p>\n<p>A niche technology is not automatically dangerous. A popular technology is not automatically safe. A good solution has a reasonable exit cost. It protects business logic, data structures, and external contracts, so that replacing it later does not damage the bones. A bad solution may be popular today, but if it mixes business, framework, storage, build, and deployment into one pot, it is mortgaging the future early.</p>\n<p>Many technical debts are borrowed under applause.</p>\n<h2>The Planning Question Has Changed</h2>\n<p>Initial planning still matters.</p>\n<p>The question has changed.</p>\n<p>The old question was: can we make the right choice the first time?</p>\n<p>Now another line has to be added: if the first choice is not fully right, can we change it later?</p>\n<p>This is closer to the real world. When many projects are born, nobody knows what they will become. Requirements change. Teams change. Traffic changes. Business models change. Ecosystems change. Asking a project to foresee its fate on day one is a little like asking a baby to fill out a retirement plan.</p>\n<p>A more reliable method is to leave passages for the future.</p>\n<p>Make interface boundaries a little clearer. Keep domain models a little cleaner. Keep tests close to real behavior. Let documentation explain key tradeoffs. Keep data migration conservative. None of this looks fashionable, but it can make the future person who has to refactor the system swear a little less.</p>\n<p>That future person is usually yourself.</p>\n<h2>Technical Debt Becomes a Ledger</h2>\n<p>Technical debt will not disappear.</p>\n<p>AI cannot eliminate laziness. It cannot eliminate complex business. It cannot eliminate wrong judgment. As long as software keeps running in the real world, debt will keep appearing. The difference is that old debt often felt like a court judgment. Once stamped, people were pinned down. Now it can feel more like a ledger. If the amount is clear, the interest is clear, and the repayment path is clear, there is room to turn things around.</p>\n<p>This matters especially for independent developers.</p>\n<p>When one person builds a project, being trapped by early choices is one of the worst outcomes. The framework is not a good fit. The ecosystem is inconvenient. The architecture gets more awkward. New features are hard to write. Old code is scary to touch. The project is not dead yet, but the developer has already been worn down by the code.</p>\n<p>AI gives small teams and independent developers a return ticket.</p>\n<p>You can start with a familiar solution and get the thing working. You can try a new framework to validate an idea. You can change the foundation while the product is still small. You can reorganize layers after the business grows into a new shape, instead of continuing to paste ointment onto the old structure.</p>\n<p>But if every impulse is called refactoring, the system will soon be shouted apart.</p>\n<p>Refactoring should move the system closer to the business itself. It should reduce the resistance of future changes. It should put scattered concepts back into place. Chasing new terms, changing skins, and adding a line to a resume are all possible activities. They do not need to borrow the name of refactoring.</p>\n<h2>Freedom Still Needs Discipline</h2>\n<p>AI has taken part of the heavy repetitive labor off developers’ shoulders.</p>\n<p>That matters.</p>\n<p>The most valuable ability of a programmer was never changing the same kind of code a thousand times. The more important ability is judgment: what deserves change, what should stay, what only works for now, and what will later tighten around the neck.</p>\n<p>With AI, refactoring becomes a little less heroic. It can be more ordinary, more frequent, and closer to what engineering should have been: observe the system, find problems, adjust the structure, verify behavior, and move on.</p>\n<p>Freedom needs discipline under it.</p>\n<p>Dare to try, and know how to clean up. Dare to begin, and dare to tear down. Do not worship the first choice. Do not treat long-term structure lightly. Design documents, tests, business boundaries, and release discipline are the foundation of this freedom.</p>\n<p>In the past, software projects often looked like trains forced onto the track of their first technical choice. If the track was crooked, the train kept moving forward, smoking all the way.</p>\n<p>In the AI era, the track is no longer so sacred.</p>\n<p>The road can change. The bridge can be rebuilt. The vehicle can be replaced. The destination must stay clear. The road needs markers. Every reroute must survive verification. If those things are in place, developers do not have to treat early choices as lifelong shackles.</p>\n<p>AI has not relieved developers of judgment.</p>\n<p>It has only lowered the wall built from repetitive labor. When the wall is lower, people can see the road beyond it.</p>\n<p>Seeing it is not enough.</p>\n<p>People still have to walk.</p>\n","date_published":"2026-05-14T00:00:00.000Z","tags":["AI","Refactoring","Architecture","Developer"],"language":"en"},{"id":"https://www.lihuanyu.com/posts/2026/AI%E6%89%93%E7%A0%B4%E4%BA%86%E4%BA%92%E8%81%94%E7%BD%91%E7%9A%84%E9%9B%B6%E8%BE%B9%E9%99%85%E6%88%90%E6%9C%AC%E7%A5%9E%E8%AF%9D/","url":"https://www.lihuanyu.com/posts/2026/AI%E6%89%93%E7%A0%B4%E4%BA%86%E4%BA%92%E8%81%94%E7%BD%91%E7%9A%84%E9%9B%B6%E8%BE%B9%E9%99%85%E6%88%90%E6%9C%AC%E7%A5%9E%E8%AF%9D/","title":"AI 打破了互联网的零边际成本神话","summary":"从 AI 应用的真实账单出发，重新理解传统互联网的低边际成本、AI 的按需生产属性，以及免费、增长和独立开发在 AI 时代为什么都要重新算账。","content_html":"<p>最近有个感觉越来越强：AI 应用不像传统互联网产品，更像是互联网后面接了一座工厂。</p>\n<p>这个说法听起来有点怪。毕竟 AI 也是网页、App、API，也是软件工程，也是订阅、会员、SaaS 这些老词。用户打开一个页面，输入一句话，得到一段回答，看起来和过去使用搜索、IM、在线工具也没什么本质区别。</p>\n<p>但账单不会骗人。</p>\n<p>传统互联网很大一部分魔法，来自复制和分发。做一个搜索引擎很贵，做一个电商平台很贵，做一个社交网络更贵。但当系统已经建起来之后，多服务一个用户的额外成本，往往低得多。</p>\n<p>这个额外成本，更准确地说叫边际成本。</p>\n<p><a href=\"/en/posts/2026/ai-broke-the-zero-marginal-cost-myth-of-the-internet/\">English version: AI Broke the Zero Marginal Cost Myth of the Internet</a></p>\n<p><img src=\"/assets/posts/2026/ai-marginal-cost/01-ai-factory-behind-internet.jpg\" alt=\"互联网后面的 AI 工厂\"></p>\n<p>边际成本低，才有了互联网行业很多后来被奉为常识的东西：免费、补贴、先增长后商业化、先把用户做起来再说。只要用户来了，后面总能想办法从广告、会员、佣金、游戏、金融、云服务或者别的什么地方把钱挣回来。</p>\n<p>这种故事听多了，人很容易产生一种错觉：软件嘛，反正复制一份又不要钱。</p>\n<p>AI 把这个错觉打碎了。</p>\n<h2>互联网以前像印刷术</h2>\n<p>传统互联网当然不是没有成本。服务器要钱，带宽要钱，工程师工资更要钱。大公司一年花在机器和人上的钱，绝不是小数目。</p>\n<p>但它的基本气质还是复制和分发。</p>\n<p>一篇文章写出来，可以被一万人看。一条商品详情页做好，可以被一万人打开。一条朋友圈发出去，可以在很多人的信息流里出现。一份搜索索引建好之后，可以服务无数次查询。哪怕背后还有缓存、数据库、推荐系统、广告系统，总体上仍然是在把已经存在的东西更高效地送到用户面前。</p>\n<p>所以互联网像印刷术。</p>\n<p>印第一本书很麻烦，排版、校对、制版、开机，都要成本。但机器一旦转起来，多印几本，单位成本就下来了。互联网把这件事做到了极致，甚至让人忘了纸张和油墨的存在。</p>\n<p>这也是为什么早期互联网公司敢烧钱。用户越多，数据越多，网络效应越强，成本被摊得越薄。增长看起来像一条通向胜利的路，虽然路上死过很多公司，但逻辑本身是通顺的。</p>\n<p>AI 不一样。AI 的很多输出不是提前印好的书，而是用户来了以后现场开炉。</p>\n<h2>AI 更像按件生产</h2>\n<p>用户问一句话，模型要推理一次。用户让它总结一篇长文，模型要读上下文再推理一次。用户让它画一张图，后面是更贵的图像模型和更长的计算时间。用户让 Agent 搜索、读文件、写代码、跑测试，那就不只是一次回答，而是一串连续的生产动作。</p>\n<p><img src=\"/assets/posts/2026/ai-marginal-cost/02-on-demand-production-line.jpg\" alt=\"AI 的按需生产线\"></p>\n<p>这时候，软件的外壳还在，工厂的性质却露出来了。</p>\n<p>token 像原材料，GPU 时间像机器工时，显存像车间容量，模型能力像设备精度，上下文长度像工艺复杂度。API 调用像找外面的工厂代工，自部署模型像自己买机器建产线。</p>\n<p>这不是一个完全严谨的经济学模型，但对开发者很有用。因为它会逼着人问一个朴素问题：</p>\n<p>每来一个用户，到底是在挣钱，还是在亏钱？</p>\n<p>过去做一个小工具，可能最担心的是服务器扛不住、数据库慢、带宽被打爆。AI 应用多了一个更扎心的问题：服务器也许扛得住，但钱包扛不住。</p>\n<p>我之前写过《<a href=\"/posts/2024/AI%E5%BA%94%E7%94%A8%E5%BC%80%E5%8F%91%E8%80%85%E7%9A%84%E5%9B%B0%E5%B1%80/\">AI 应用开发者的困局：用户来了，账单也来了</a>》。里面提到过一个 AI 绘图小程序，即使用了弹性部署，只在有用户请求时才启动 GPU，按秒计费，一张图也要 1 到 2 角钱。</p>\n<p>一两角钱，听起来很便宜。</p>\n<p>但如果这是一个免费功能，就完全是另一回事了。用户生成一张图，开发者掏一两角。用户生成十张图，开发者掏一两块。用户觉得“这功能真好玩”，开发者看账单觉得“这事不太妙”。</p>\n<p>这就是 AI 应用和普通互联网工具最不一样的地方。它不是多几个访问、多几次点击那么简单。它的每一次有效使用，都可能是真金白银的消耗。</p>\n<h2>免费开始变得沉重</h2>\n<p>互联网产品喜欢免费。免费邮箱、免费网盘、免费社交、免费内容、免费工具，都是这套逻辑养出来的。</p>\n<p><img src=\"/assets/posts/2026/ai-marginal-cost/03-free-meter.jpg\" alt=\"免费入口背后仍然有会转动的机器\"></p>\n<p>当然，免费从来不是真的免费。有人付广告费，有人付会员费，有人贡献数据，有人被生态绑定。只是用户感知不到，或者暂时不用直接掏钱。</p>\n<p>AI 应用的免费比较尴尬。</p>\n<p>用户不是只占了一个账号、几行数据库记录、几 MB 存储空间。用户只要开始真正使用模型，就开始烧成本。长上下文、多轮对话、图片生成、语音生成、视频生成、联网搜索、代码执行，这些东西越强，越不像空气。</p>\n<p>于是很多过去可以后置的问题，现在必须提前想清楚。</p>\n<p>未登录用户能不能用？免费额度给多少？高成本模型要不要限制？失败重试算不算额度？接口被刷怎么办？用户只是来玩一下就走，成本算谁的？付费用户的收入能不能覆盖模型账单？</p>\n<p>这些问题看起来是商业问题，落到代码里全是工程问题。</p>\n<p>要做登录，要做额度，要做限流，要做缓存，要做队列，要做模型分级，要看调用成本，要防止有人把接口当公共水龙头。以前做个网页工具，裸奔上线也不是不能活。AI 工具裸奔上线，有时像把一台开着的机器放在路边，还贴一张纸：欢迎免费使用。</p>\n<p>当然会有人来用。</p>\n<p>问题是机器归谁供电。</p>\n<h2>增长也可能是一种危险</h2>\n<p>互联网人通常怕没人用。</p>\n<p>AI 产品当然也怕没人用，但它还怕另一件事：太多人来用，而且都是不付钱的人。</p>\n<p><img src=\"/assets/posts/2026/ai-marginal-cost/04-growth-cost-balance.jpg\" alt=\"增长和成本需要一起看\"></p>\n<p>这件事对独立开发者尤其残酷。大公司可以把 AI 当战略投入，可以用云、广告、生态、资本市场来摊账。独立开发者没有这么多腾挪空间。账单来了就是账单，信用卡扣款不会因为“这是未来趋势”就少扣一点。</p>\n<p>所以 AI 应用的增长要看质量。</p>\n<p>一个愿意为工作流效率付费的用户，和一个生成几张图就走的用户，对产品的意义完全不同。前者可能是业务，后者可能只是成本。以前说用户增长，多少还有点喜气；AI 应用里，低质量增长有时像接到一批没有货款的订单，车间加班加点，老板越忙越穷。</p>\n<p>这也是为什么 AI 应用更早进入毛利思维。</p>\n<p>一个功能酷不酷，不够。用户想不想用，也不够。还要看用得越多亏不亏。很多 AI demo 在展示时很惊艳，真正放到线上就会露出另一面：效果越好，大家越爱用；大家越爱用，账单越好看。</p>\n<p>当然，是反方向的好看。</p>\n<h2>成本会降，但不会消失</h2>\n<p>有人会说，算力会越来越便宜，模型会越来越便宜。</p>\n<p>我也相信会便宜。芯片会进步，推理框架会优化，模型会量化、蒸馏、裁剪，小模型会承担更多简单任务，大模型厂商也会继续打价格战。</p>\n<p>但便宜不等于没有成本。</p>\n<p>宽带变便宜以后，互联网没有停留在文字网页，而是走向图片、视频、直播、云游戏。存储变便宜以后，人们也没有少存东西，而是拍更多照片、传更多视频、做更多备份。</p>\n<p>算力变便宜以后，AI 大概率也不会停留在今天这种问答强度。上下文会更长，Agent 会更复杂，自动化任务会更多，原本一天调用几次的东西，可能变成后台持续运行。成本下降会扩大使用边界，也会制造新的消耗方式。</p>\n<p>所以真正需要的不是幻想成本归零，而是学会算账。</p>\n<p>简单问题用便宜模型，复杂问题再用强模型。能缓存就缓存，能异步就异步，能让用户确认就不要重复生成，能在本地做的预处理不要全丢给大模型。模型路由、成本监控、额度体系、失败重试策略，这些听起来像工程细节，其实都是 AI 产品的生意基础。</p>\n<p>一个不看成本的 AI 应用，就像一个不看电表的工厂。机器声越响，未必越兴旺。</p>\n<h2>旧互联网公式不够用了</h2>\n<p>AI 当然仍然是软件。</p>\n<p>它可以快速迭代，可以在线分发，可以订阅收费，可以用很小的团队做出过去很难想象的东西。这些都是软件的优势。</p>\n<p>但 AI 又不只是软件。</p>\n<p>传统软件最厉害的是复制成本低。传统互联网最厉害的是分发成本低。AI 应用多了一个麻烦：高质量输出有生产成本，而且这个成本会随着使用量一起增长。</p>\n<p>所以“先免费做大规模，再慢慢商业化”这句话，在 AI 应用里要重新掂量。</p>\n<p>成本由谁承担？</p>\n<p>用户直接付费，那产品就要值得付费。企业客户买单，那就要嵌进真实工作流。广告覆盖成本，那流量价值要足够高。平台补贴，那就要接受平台什么时候想补、什么时候不想补。如果只是开发者自己承担，那最好一开始就知道，这是练手、实验，还是一门长期生意。</p>\n<p>不是所有 AI 项目都要赚钱。学习项目、作品集、技术验证，当然可以不赚钱。但如果把它当产品，就不能只讲愿景，不看账本。</p>\n<p>互联网过去让人相信，规模会解决很多问题。</p>\n<p>AI 会提醒人们，规模也会放大很多问题。</p>\n<h2>最后还是那句老话</h2>\n<p>所以，“AI 像制造业”不是一个单纯的比喻。</p>\n<p>它是在提醒开发者：AI 把生产行为重新放回了每一次请求里。传统互联网像复制和分发，AI 更像按需生产。它仍然有软件的速度，却也有工厂的成本；它可以把入口开给全世界，也会在每一次输出时消耗真实资源。</p>\n<p>这并不悲观。</p>\n<p>相反，正因为成本真实，价值也会更真实。一个 AI 应用如果能让用户愿意为每一次生产付费，或者愿意为它带来的效率长期付费，那它就不只是玩具。它可能是工具，可能是服务，也可能是一种新的生产组织方式。</p>\n<p>只是这个时代不再允许开发者完全躲在“互联网免费”的幻觉里。</p>\n<p>软件把复制变得便宜，互联网把分发变得便宜，AI 则让计算本身变成产品的一部分。</p>\n<p>而生产这件事，说到底从来不神秘：</p>\n<p>机器要转，材料要耗，电表要走。有人愿意为它付钱，生意才可能继续。</p>\n","date_published":"2026-05-11T00:00:00.000Z","tags":["AI","互联网","商业模式","开发者"],"language":"zh"},{"id":"https://www.lihuanyu.com/en/posts/2026/ai-broke-the-zero-marginal-cost-myth-of-the-internet/","url":"https://www.lihuanyu.com/en/posts/2026/ai-broke-the-zero-marginal-cost-myth-of-the-internet/","title":"Why AI Has Higher Marginal Costs Than Internet Software","summary":"Marginal cost is the cost of serving one more unit. AI adds inference, tokens, GPU time, and tool calls to each request, changing software unit economics.","content_html":"<p>Marginal cost is the additional cost of producing or serving one more unit. Traditional internet software can distribute an existing page, file, or database result again. An artificial intelligence (AI) product often performs new inference for every answer, image, or agent run.</p>\n<p>This difference does not make traditional software free or AI unprofitable. It means usage and cost are more tightly coupled in AI products, so developers must consider unit economics earlier.</p>\n<p><a href=\"/posts/2026/AI%E6%89%93%E7%A0%B4%E4%BA%86%E4%BA%92%E8%81%94%E7%BD%91%E7%9A%84%E9%9B%B6%E8%BE%B9%E9%99%85%E6%88%90%E6%9C%AC%E7%A5%9E%E8%AF%9D/\">Chinese version of this article</a></p>\n<h2>Compare marginal cost in internet software and AI</h2>\n<p>Traditional internet products spend heavily on engineering, infrastructure, moderation, and operations. After producing a piece of content or computing a reusable result, however, another request may need only cache lookup, request processing, and bandwidth.</p>\n<p>AI retains those software costs and adds per-request production work:</p>\n<div class=\"table-scroll\"><table>\n<thead>\n<tr>\n<th>Cost dimension</th>\n<th>Traditional internet software</th>\n<th>AI application</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Primary request work</td>\n<td>Retrieve or transform stored data</td>\n<td>Run model inference and possibly external tools</td>\n</tr>\n<tr>\n<td>Reuse</td>\n<td>Cache pages, files, and computed results</td>\n<td>Cache repeated inputs, but many prompts require new output</td>\n</tr>\n<tr>\n<td>Variable cost drivers</td>\n<td>Bandwidth, storage operations, database load</td>\n<td>Tokens, GPU time, model choice, context size, and tool calls</td>\n</tr>\n<tr>\n<td>Growth risk</td>\n<td>Capacity and infrastructure scale with traffic</td>\n<td>Infrastructure and model spend scale with actual usage</td>\n</tr>\n</tbody>\n</table>\n</div><p>The comparison is about cost structure, not absolute cost. A database-heavy service can be expensive, while a small local model can be cheap. The key question is whether one more successful task triggers new production work.</p>\n<p><img src=\"/assets/posts/2026/ai-marginal-cost/01-ai-factory-behind-internet.jpg\" alt=\"An AI factory behind the internet\"></p>\n<p>Low marginal cost is what made many internet habits feel natural: free products, subsidies, growth before monetization, “get users first and figure out the business later.” Once users arrive, there is always some story about ads, memberships, commissions, games, finance, cloud services, or something else that can pay the bill later.</p>\n<p>Hear that story often enough, and it creates an illusion: software is basically free to copy.</p>\n<p>AI breaks that illusion.</p>\n<h2>Traditional internet software resembles printing</h2>\n<p>Traditional internet products were never actually free to run.</p>\n<p>Servers cost money. Bandwidth costs money. Storage costs money. Engineers cost much more. At large scale, the infrastructure bill of an internet company is not some rounding error.</p>\n<p>But the basic character of the internet was still copying and distribution.</p>\n<p>An article can be written once and read by ten thousand people. A product detail page can be built once and opened by ten thousand shoppers. A social post can enter many feeds. A search index, once built, can serve countless queries. There are still caches, databases, recommendation systems, ad systems, and moderation systems behind it all, but the broad pattern is the same: take something that already exists and deliver it more efficiently.</p>\n<p>In that sense, the internet was closer to printing.</p>\n<p>The first copy of a book is hard. Editing, layout, plates, machines, logistics, all of it costs money. But once the machine is running, printing more copies brings the unit cost down. The internet pushed this logic so far that people almost forgot the paper and ink existed.</p>\n<p>That is why early internet companies could burn money with some internal logic. More users meant more data, stronger network effects, and costs spread across a larger base. Growth looked like a road toward victory. Many companies died on that road, of course, but the logic itself was coherent.</p>\n<p>AI is different. Much of what AI produces is not a book printed in advance. It is more like firing up the furnace after the user arrives.</p>\n<h2>AI resembles on-demand production</h2>\n<p>A user asks a question, and the model runs inference. A user asks for a long document summary, and the model reads context and runs inference again. A user asks for an image, and a more expensive image model may run for longer. A user starts an agent task that searches, reads files, writes code, and runs tests, and the cost becomes a chain of production steps.</p>\n<p><img src=\"/assets/posts/2026/ai-marginal-cost/02-on-demand-production-line.jpg\" alt=\"AI as on-demand production\"></p>\n<p>The software shell is still there, but the factory starts showing through.</p>\n<p>Tokens are raw materials. GPU time is machine time. VRAM is workshop capacity. Model quality is equipment precision. Context length is process complexity. API calls are outsourced manufacturing. Self-hosting a model is buying machines and building your own line.</p>\n<p>This is not a perfect economic model, but it is a useful one for developers, because it forces a plain question:</p>\n<p>When one more user arrives, are you making money or losing money?</p>\n<p>In the past, a small online tool mostly worried about whether the server could handle traffic, whether the database was slow, or whether bandwidth would spike. AI applications add a sharper question: the server may survive, but will the wallet survive?</p>\n<p>I wrote about this before in <a href=\"/posts/2024/AI%E5%BA%94%E7%94%A8%E5%BC%80%E5%8F%91%E8%80%85%E7%9A%84%E5%9B%B0%E5%B1%80/\">The Dilemma of AI Application Developers</a>. One example was an AI image generation mini program I built. Even with elastic deployment, starting the GPU only when there was a user request and charging by the second, one generated image still cost about 0.1 to 0.2 RMB.</p>\n<p>That sounds cheap for one image.</p>\n<p>But if the feature is free, the meaning changes completely. A user generates one image, and the developer pays a little. A user generates ten images, and the developer pays more. The user thinks, “This is fun.” The developer looks at the bill and thinks, “This is not going well.”</p>\n<p>That is the difference between many AI tools and ordinary internet tools. It is not just a few more page views or clicks. Every real use can become a real cost.</p>\n<h2>Free usage creates a direct variable cost</h2>\n<p>Internet products love being free.</p>\n<p>Free email, free cloud storage, free social networks, free content, free utilities. None of them were truly free, of course. Someone paid through ads, memberships, data, ecosystem lock-in, or some delayed business model. Users just did not feel the cost directly, at least not at the beginning.</p>\n<p><img src=\"/assets/posts/2026/ai-marginal-cost/03-free-meter.jpg\" alt=\"A free entrance can still lead to a running machine\"></p>\n<p>Free AI applications are more awkward.</p>\n<p>A user is not merely taking up an account, a few database rows, or some storage. Once the user actually uses the model, cost starts burning. Long context, multi-turn conversations, image generation, speech generation, video generation, web search, code execution: the stronger these features become, the less they resemble air.</p>\n<p>So questions that used to be postponed now have to be answered early.</p>\n<p>Can anonymous users use it? How much free quota should there be? Should expensive models be restricted? Do failed retries count against quota? What happens if the API is abused? If a user only plays with it once and leaves, who pays for that? Can the revenue from paid users cover the model bill?</p>\n<p>These look like business questions. In code, they are engineering questions.</p>\n<p>You need login. You need quota. You need rate limits. You need caching. You need queues. You need model tiers. You need cost monitoring. You need protection against people treating your API like a public tap. In the old days, a small web utility could sometimes launch in a fairly naked state and survive. Launching a naked AI utility is more like putting a running machine on the street with a note that says: free to use.</p>\n<p>People will use it.</p>\n<p>The question is who pays for the electricity.</p>\n<h2>Growth can increase losses</h2>\n<p>Internet people are usually afraid that nobody will use their product.</p>\n<p>AI products are afraid of that too. But they are also afraid of something else: too many people using the product without paying.</p>\n<p><img src=\"/assets/posts/2026/ai-marginal-cost/04-growth-cost-balance.jpg\" alt=\"Growth and cost need to be measured together\"></p>\n<p>This is especially harsh for independent developers. Large companies can treat AI as a strategic investment. They can spread the cost across cloud businesses, ads, ecosystems, financing, and long-term positioning. Independent developers do not have that much room to maneuver. A bill is a bill. The credit card charge does not shrink because “AI is the future.”</p>\n<p>So the quality of growth matters more.</p>\n<p>A user willing to pay for workflow efficiency and a user who generates a few images and disappears mean very different things to a product. The first may be a business. The second may only be cost. In the old internet, user growth at least sounded cheerful. In AI, low-quality growth can feel like receiving a pile of orders with no payment attached. The workshop is busy, the owner is poorer.</p>\n<p>That is why AI products enter gross margin thinking earlier.</p>\n<p>It is not enough for a feature to be cool. It is not enough for users to want it. You also have to ask whether more usage makes the product lose more money. Many AI demos are impressive in a presentation and much less comforting online. The better the effect, the more people use it. The more people use it, the more beautiful the bill becomes.</p>\n<p>Beautiful in the wrong direction.</p>\n<h2>Falling inference prices do not remove unit economics</h2>\n<p>One common reply is that compute will get cheaper and models will get cheaper.</p>\n<p>I believe that too. Chips will improve. Inference frameworks will get faster. Models will be quantized, distilled, routed, and specialized. Smaller models will handle more simple tasks. Large model providers will keep fighting on price.</p>\n<p>But cheaper is not the same as free.</p>\n<p>When bandwidth became cheaper, the internet did not stay with text pages. It moved to images, video, livestreaming, and cloud gaming. When storage became cheaper, people did not store less. They took more photos, uploaded more videos, and kept more backups.</p>\n<p>When compute becomes cheaper, AI will probably not stay at today’s level of usage. Context windows will grow. Agents will become more complex. Automated tasks will run more often. Something that is called a few times a day may become something that runs continuously in the background. Falling cost expands the boundary of use, but it also creates new ways to consume resources.</p>\n<p>So the real answer is not to wait for cost to become zero. The real answer is to learn how to account for it.</p>\n<p>Use cheap models for simple tasks and stronger models only when needed. Cache what can be cached. Run what can be asynchronous outside the realtime path. Ask for confirmation instead of regenerating blindly. Do local preprocessing when possible instead of sending everything to a large model. Model routing, cost monitoring, quota design, retry policy: these sound like engineering details, but they are business fundamentals for AI products.</p>\n<p>An AI application that does not watch cost is like a factory that does not watch the power meter. Loud machines do not necessarily mean a healthy business.</p>\n<h2>The old internet growth formula is not enough</h2>\n<p>AI is still software.</p>\n<p>It can iterate quickly. It can be distributed online. It can be sold by subscription. A small team can build things that would have been hard to imagine before. These are real software advantages.</p>\n<p>But AI is not only software.</p>\n<p>Traditional software was powerful because copying was cheap. The internet was powerful because distribution was cheap. AI applications add a difficult layer: high-quality output has production cost, and that cost rises with usage.</p>\n<p>So the old phrase “grow first, monetize later” needs to be weighed again.</p>\n<p>Who pays for each service?</p>\n<p>If users pay directly, the product must be worth paying for. If enterprises pay, the product must enter real workflows. If ads cover the cost, traffic value must be high enough. If a platform subsidizes the cost, the product must accept that the platform can change its mind. If the developer pays personally, it is better to know from the beginning whether this is a learning project, an experiment, or a long-term business.</p>\n<p>Not every AI project needs to make money. Learning projects, demos, portfolios, and technical experiments can have their own value. But if something is treated as a product, it cannot live only on vision while ignoring the ledger.</p>\n<p>The old internet taught people that scale solves many problems.</p>\n<p>AI reminds people that scale also amplifies many problems.</p>\n<h2>Treat model usage as production cost</h2>\n<p>So “AI is like manufacturing” is not just a colorful metaphor.</p>\n<p>It is a reminder that AI brings production back into each request. Traditional internet products were closer to copying and distribution. AI is closer to on-demand production. It still has the speed of software, but it also has the cost discipline of a factory. It can open the entrance to the whole world, and it can spend real resources on every output.</p>\n<p>This is not pessimistic.</p>\n<p>In fact, because the cost is real, the value can become more real too. If an AI application can make users willing to pay for each act of production, or pay continuously for the efficiency it creates, then it is not just a toy. It may be a tool, a service, or a new way to organize work.</p>\n<p>But this era no longer lets developers hide completely inside the old illusion of free internet products.</p>\n<p>Software made copying cheap. The internet made distribution cheap. AI makes computation itself part of the product.</p>\n<p>And production has always had a simple rule:</p>\n<p>Machines run. Materials are consumed. The meter moves. Someone has to pay, or the business cannot continue.</p>\n","date_published":"2026-05-11T00:00:00.000Z","date_modified":"2026-08-01T00:00:00.000Z","tags":["AI","Internet","Business Model","Developer"],"language":"en"},{"id":"https://www.lihuanyu.com/en/posts/2026/i-built-a-fire-calculator-financial-freedom-should-not-depend-on-vibes/","url":"https://www.lihuanyu.com/en/posts/2026/i-built-a-fire-calculator-financial-freedom-should-not-depend-on-vibes/","title":"I Built a FIRE Calculator Because Financial Freedom Should Not Depend on Vibes","summary":"A personal reflection on why I built ChooseFIRE, and how a simple calculator can make income, spending, savings rate, returns, and personal freedom easier to reason about.","content_html":"<p>When I first came across FIRE, the part that caught my attention was early retirement.</p>\n<p>It is an attractive idea. When work feels intense and life is pushed forward by meetings, messages, and deadlines, it is easy to picture FIRE as a clean finish line: save enough money, leave the workplace, and never wake up to an alarm again.</p>\n<p>My understanding changed over time. What attracts me now is the extra room people can have when making life decisions. The ability to leave a bad work environment. The ability to avoid making every career decision under short-term cash pressure. The ability to spend more time on things that actually matter.</p>\n<p>Financial freedom sounds like a big phrase, but in real life it often turns into a few concrete questions. How much do I spend each year? How much do I already have? How much can I save every month? What happens if returns are lower? How much does the target change if my annual spending changes?</p>\n<p>Those questions are simple on paper. They are still hard to answer by feeling alone. That is why I built a small tool: <a href=\"https://choosefire.com/\">ChooseFIRE</a>.</p>\n<p><a href=\"/posts/2026/%E6%88%91%E5%81%9A%E4%BA%86%E4%B8%80%E4%B8%AAFIRE%E8%AE%A1%E7%AE%97%E5%99%A8-%E8%B4%A2%E5%8A%A1%E8%87%AA%E7%94%B1%E4%B8%8D%E8%AF%A5%E5%8F%AA%E9%9D%A0%E6%84%9F%E8%A7%89/\">Chinese version of this article</a></p>\n<h2>Knowing the Concept Is Different From Knowing Where You Are</h2>\n<p>Several ideas show up repeatedly in FIRE discussions.</p>\n<p>The 4% rule is probably the most common one. In simple terms, it says that if your assets reach about 25 times your annual spending, a relatively low withdrawal rate may be enough to cover your living costs. Savings rate, passive income, Lean FIRE, Fat FIRE, and Coast FIRE all circle around the same broad question: how can assets gradually take over the cost of living?</p>\n<p>These ideas are useful. They show that financial independence has a structure. It is not pure fantasy, and it is not reserved only for people with extremely high income. There are variables that can be discussed.</p>\n<p>But once you try to apply them to your own life, the uncertainty appears quickly.</p>\n<p>Should annual spending be based on your current lifestyle or your expected post-retirement lifestyle? What return assumption is too optimistic? How should inflation be treated? If you save a little more every month, how much does it really change the timeline? If you move to a different city, does the target change immediately?</p>\n<p>An article can introduce the concepts, but it cannot answer these questions for everyone. Income structure, family responsibilities, city, spending habits, and risk tolerance are all personal. When someone reads “25 times annual spending,” the missing step is usually not the formula. It is the process of putting that formula into their own life.</p>\n<h2>Why I Built ChooseFIRE</h2>\n<p>The immediate reason for building ChooseFIRE was simple: I wanted FIRE calculations to feel more visible.</p>\n<p>Many calculators ask for a few numbers and return a result. That result can be useful, but the most interesting part of FIRE is often not the final number. It is what happens when you adjust the inputs.</p>\n<p>What if annual spending increases by 20%? What if monthly savings increase by $300? What if expected return drops from 6% to 4%? Does the plan still make sense? These changes are often more informative than a single answer such as “you need 17 more years.”</p>\n<p>So ChooseFIRE does not try to produce a dramatic final verdict. It puts the key variables on the same page:</p>\n<ul>\n<li>Current assets</li>\n<li>Annual spending</li>\n<li>Regular savings</li>\n<li>Expected return</li>\n<li>Target withdrawal rate</li>\n<li>Time to reach the target</li>\n</ul>\n<p>Once these numbers are visible together, some things become clearer. Some people may find that their goal is closer than they imagined. Others may find that the real bottleneck is not income, but a spending structure they have never seriously examined.</p>\n<p>I did not want to build a complicated personal finance product. I wanted something closer to an editable worksheet: write down the current situation, change the assumptions, and see where different choices lead.</p>\n<h2>Spending Is Easy to Underestimate</h2>\n<p>Spending has a special role in FIRE calculations.</p>\n<p>Higher income certainly helps. Higher investment returns make compounding more powerful. But income and returns are not fully stable. Income depends on industry, company, cycle, and location. Returns are even less controllable; nobody can lock in long-term market performance in advance.</p>\n<p>Spending is not fully controllable either. Rent, mortgages, medical costs, education, and family responsibilities cannot be solved by saying “just spend less.” Still, compared with investment returns, spending is often closer to lifestyle design and long-term personal choices.</p>\n<p>Take a simple example.</p>\n<p>If someone spends $60,000 per year, a 4% withdrawal rate implies a target of about $1.5 million. If annual spending drops to $45,000, the target becomes about $1.125 million. On the yearly budget, that difference is $15,000. In a FIRE target, it becomes $375,000.</p>\n<p>This is why savings rate matters so much in FIRE discussions. A higher savings rate works in two directions at the same time. You invest more each year, and if the higher savings rate comes from lower spending, the final required portfolio also becomes smaller.</p>\n<p>This does not mean everyone should live as cheaply as possible. Quality of life, health, relationships, and long-term happiness should not be flattened into a spreadsheet. What I care about is understanding the long-term cost of choices. Once the cost is visible, the decision becomes more honest.</p>\n<h2>A Calculator Cannot Make Life Decisions</h2>\n<p>I do not want ChooseFIRE to be treated as a tool that tells people what to do.</p>\n<p>It does not tell you whether you should retire. It does not tell you what assets to buy. It cannot guarantee any future return. The 4% rule itself is only a common historical framework, and it needs to be interpreted carefully across countries, tax systems, inflation environments, portfolio choices, and personal risk tolerance.</p>\n<p>Calculation still has value. It can turn questions that feel emotional into numbers that can be discussed.</p>\n<p>Someone may feel that financial independence is forever impossible, then find that the main issue is a low current savings rate. Someone else may feel almost ready to stop working, then discover that the plan becomes fragile if returns are two percentage points lower.</p>\n<p>Neither result is the final answer. Both results make the risk more visible.</p>\n<p>Personal finance is hard because it is practical and emotional at the same time. Anxiety can make the goal feel unreachable. Optimism can make uncertainty look smaller than it is. A rough but transparent calculation can bring the discussion back to variables that can be adjusted.</p>\n<h2>FIRE Is More Than the Day You Quit</h2>\n<p>After building this tool, I have come to see financial freedom more as a spectrum.</p>\n<p>Fully covering all living expenses is one state, but there are many meaningful states before that.</p>\n<p>Having enough savings for six months of expenses already makes unemployment less frightening. Having assets that can cover several years of living costs makes it easier to change jobs, switch fields, or take a break. At a certain point, even before full FIRE, a person may reach something closer to Coast FIRE or Barista FIRE: the need to keep aggressively accumulating capital becomes lower, and only part of the cash flow has to be covered by work.</p>\n<p>These middle states are less dramatic than “early retirement,” but they are closer to real life.</p>\n<p>Most people do not suddenly jump from full-time work to permanent retirement. More often, they gradually gain options. They can say no to unreasonable work. They can choose work that pays less but fits better. They can leave more space for family and health. They can slowly build a personal project before it has to pay the bills.</p>\n<p>That is the feeling I want ChooseFIRE to support. Financial freedom can be a final portfolio number, but it can also be a way to understand your current position and the choices around it.</p>\n<h2>Start With One Number</h2>\n<p>If you are interested in FIRE, I do not think the first question has to be “When can I retire?”</p>\n<p>A better starting point may be three numbers:</p>\n<ul>\n<li>How much do I spend each year now?</li>\n<li>How much would I need to cover that spending?</li>\n<li>At my current savings pace, how far away is that target?</li>\n</ul>\n<p>Those questions are already enough to reduce a lot of uncertainty.</p>\n<p>Then you can start changing the assumptions. What happens if spending rises? What happens if it falls? What if expected returns are more conservative? What if monthly savings increase a little? After a few rounds, FIRE stops being a vague internet concept and becomes a set of tradeoffs connected to your own life.</p>\n<p>That is why I built <a href=\"https://choosefire.com/\">ChooseFIRE</a>.</p>\n<p>It cannot replace investment judgment or life decisions. But if it helps someone seriously look at the relationship between income, spending, assets, and time for the first time, then it is doing something useful.</p>\n<p>Financial freedom should not depend on vibes. At the very least, put the numbers on the table first.</p>\n","date_published":"2026-05-08T00:00:00.000Z","tags":["FIRE","Financial Independence","Early Retirement","Personal Project"],"language":"en"},{"id":"https://www.lihuanyu.com/posts/2026/%E6%88%91%E5%81%9A%E4%BA%86%E4%B8%80%E4%B8%AAFIRE%E8%AE%A1%E7%AE%97%E5%99%A8-%E8%B4%A2%E5%8A%A1%E8%87%AA%E7%94%B1%E4%B8%8D%E8%AF%A5%E5%8F%AA%E9%9D%A0%E6%84%9F%E8%A7%89/","url":"https://www.lihuanyu.com/posts/2026/%E6%88%91%E5%81%9A%E4%BA%86%E4%B8%80%E4%B8%AAFIRE%E8%AE%A1%E7%AE%97%E5%99%A8-%E8%B4%A2%E5%8A%A1%E8%87%AA%E7%94%B1%E4%B8%8D%E8%AF%A5%E5%8F%AA%E9%9D%A0%E6%84%9F%E8%A7%89/","title":"我做了一个 FIRE 计算器：财务自由不该只靠感觉","summary":"从接触 FIRE 的个人感受出发，聊聊为什么做 ChooseFIRE，以及一个计算器能如何帮助人更清楚地理解收入、支出、储蓄率和选择权之间的关系。","content_html":"<p>第一次接触 FIRE 的时候，我和很多人一样，最先注意到的是“提前退休”。</p>\n<p>这四个字很有吸引力。尤其是在工作压力比较大、生活节奏被会议和消息推着走的时候，很容易把 FIRE 想象成一个终点：攒够一笔钱，然后离开职场，再也不用被闹钟叫醒。</p>\n<p>后来我对这件事的理解慢慢变了。真正吸引我的，是人在面对生活选择时能多一点余地。比如不必为了现金流忍受糟糕的工作环境，不必在职业低谷时被短期收入逼着做决定，也可以把一部分时间投向自己真正想做的事情。</p>\n<p>财务自由听起来像一个很大的词，落到个人身上，其实经常只是一些具体问题：每年要花多少钱？现在有多少资产？每个月还能存下多少？如果收益率低一点，时间会拉长多少？如果支出少一点，目标本金会少多少？</p>\n<p>这些问题不算复杂，但只靠感觉很难想清楚。所以我做了一个小工具：<a href=\"https://choosefire.com/\">ChooseFIRE</a>。</p>\n<p><a href=\"/en/posts/2026/i-built-a-fire-calculator-financial-freedom-should-not-depend-on-vibes/\">English version: I Built a FIRE Calculator Because Financial Freedom Should Not Depend on Vibes</a></p>\n<p><img src=\"/assets/posts/2026/fire-calculator/01-financial-freedom-map.jpg\" alt=\"把模糊的财务自由问题变成可见的路径\"></p>\n<h2>知道概念，和知道自己在哪，是两回事</h2>\n<p>FIRE 圈子里有几个很常见的概念。</p>\n<p>比如 4% 法则。它大致表达的是，如果一个人的资产达到年支出的 25 倍，理论上就可以通过较低比例的年度提取来覆盖生活开销。还有储蓄率、被动收入、Lean FIRE、Fat FIRE、Coast FIRE 这些说法，也都围绕同一个问题展开：怎样让资产逐渐承担生活成本。</p>\n<p><img src=\"/assets/posts/2026/fire-calculator/03-target-multiple.jpg\" alt=\"用 25 倍年支出理解 FIRE 目标\"></p>\n<p>这些概念有帮助。它们至少让人意识到，财务自由并不完全是玄学，也不是只属于少数高收入人群的想象。它背后有一组可以讨论的变量。</p>\n<p>但真正轮到自己算的时候，模糊感很快就出现了。</p>\n<p>年支出到底应该按现在的水平算，还是按退休后的水平算？投资收益率应该假设多少才不算太乐观？通胀要不要考虑？现在多存一点钱，对最终时间线到底有多大影响？如果换一个城市生活，目标会不会立刻变得不一样？</p>\n<p>这些问题只看一篇文章很难得到答案。因为每个人的收入结构、家庭责任、城市、消费习惯和风险承受能力都不一样。一个人在网上看到“25 倍年支出”时，更需要的是把公式代入自己生活的过程。</p>\n<h2>为什么想做 ChooseFIRE</h2>\n<p>做 ChooseFIRE 的直接原因，是我希望 FIRE 测算能更直观一点。</p>\n<p>很多计算器会让人输入几个数字，然后给出一个结果。这个结果当然有用，但 FIRE 真正有意思的地方，往往不在那个最终数字，而在调整变量时发生的变化。</p>\n<p>年支出增加 20%，目标本金会怎么变？每个月多存 2000 元，时间会提前多少？投资收益率从 6% 调到 4%，计划还站得住吗？这些变化比单独一个“你还需要多少年”的答案更有价值。</p>\n<p>所以 ChooseFIRE 没有把重点放在一个确定命运的数字上。我更希望它能帮人把几个关键变量摆到桌面上：</p>\n<ul>\n<li>当前资产</li>\n<li>年支出</li>\n<li>定期储蓄</li>\n<li>预期收益率</li>\n<li>目标提取率</li>\n<li>距离目标的时间</li>\n</ul>\n<p><img src=\"/assets/posts/2026/fire-calculator/02-variables-table.jpg\" alt=\"把关键变量放在同一个页面里\"></p>\n<p>这些数字一旦放在同一个页面里，很多事情会变得清楚。比如有些人会发现，自己离目标并没有想象中那么远；也有些人会发现，真正拖慢进度的不是收入，而是支出结构一直没有被认真看过。</p>\n<p>我不想把它做成一个复杂的理财产品。它更像是一张可以反复改的草稿纸：先把当前状态写下来，再看看不同选择会把自己带到哪里。</p>\n<h2>支出是最容易被低估的变量</h2>\n<p>在 FIRE 测算里，支出有一个很特别的位置。</p>\n<p>收入越高，当然越容易积累资产。投资收益率越高，复利也越明显。但收入和收益率都没有那么稳定。收入受行业、公司、周期、城市影响很大；收益率更不用说，长期市场表现没人能提前锁定。</p>\n<p>支出也不是完全可控。房租、房贷、医疗、教育、家庭责任，这些都不是一句“少花点”就能解决的。但相比收益率，支出至少更接近个人生活方式和长期选择。</p>\n<p><img src=\"/assets/posts/2026/fire-calculator/04-spending-lever.jpg\" alt=\"支出变化会改变通往目标的路径\"></p>\n<p>举个简单例子。</p>\n<p>如果一个人每年支出 20 万，按 4% 提取率估算，需要大约 500 万资产来覆盖这部分支出。如果年支出降到 15 万，目标资产就变成 375 万。账面上一年少花 5 万，映射到 FIRE 目标里，会少掉 125 万目标本金。</p>\n<p>这也是 FIRE 讨论里经常强调储蓄率的原因。储蓄率提高带来的影响有两层：一方面每年能投入更多资金，另一方面如果支出下降，最终需要的本金也会下降。两边同时变化时，时间线可能会明显缩短。</p>\n<p>当然，这不代表每个人都应该极端节省。生活质量、健康、家庭关系、长期幸福感，都不应该被一个表格压扁。我更看重的是看清不同消费选择背后的长期影响。知道一件事的代价之后，再决定要不要为它付钱，这会更踏实。</p>\n<h2>计算器不能替代人生决定</h2>\n<p>我不希望 ChooseFIRE 被理解成一个给人下判断的工具。</p>\n<p>它不会告诉你该不该退休，也不会告诉你应该买什么资产，更不能保证某个收益率一定会实现。4% 法则本身也只是一个常见的历史经验框架，放到不同国家、税制、通胀环境、资产配置和个人风险偏好下，都需要重新理解。</p>\n<p>但计算仍然有价值。它能把一些原本混在情绪里的问题变成可以讨论的数字。</p>\n<p>比如一个人觉得自己“永远不可能财务自由”，算完之后也许会发现，问题集中在当前储蓄率太低。另一个人觉得自己“差不多可以停下来了”，算完之后也许会发现，只要收益率低两个点，计划就会变得很脆弱。</p>\n<p>这些结果都不是最终答案，却能让人更清楚地看到风险在哪里。</p>\n<p>很多个人财务问题最麻烦的地方，是它们既现实，又容易被情绪放大。焦虑的时候会觉得永远不够，乐观的时候又容易低估未来不确定性。一个粗略但透明的测算，至少能让讨论回到可调整的变量上。</p>\n<h2>FIRE 不止是辞职那一天</h2>\n<p>做这个工具之后，我越来越觉得财务自由更像一个连续光谱。</p>\n<p><img src=\"/assets/posts/2026/fire-calculator/05-choice-spectrum.jpg\" alt=\"财务自由更像一段连续的选择光谱\"></p>\n<p>完全覆盖所有生活支出当然是一种状态，但在它之前，还有很多中间状态同样重要。</p>\n<p>有一笔能覆盖半年生活的储蓄，已经能让人面对失业时少一点慌张。有一笔能覆盖几年生活的资产，换工作、转行业、休整一段时间都会更从容。如果资产增长到一定阶段，即使还没有完全 FIRE，也可能进入 Coast FIRE 或 Barista FIRE 这样的状态：不再需要像过去那样拼命积累本金，只需要维持一部分现金流。</p>\n<p>这些中间状态不如“提前退休”醒目，但它们更贴近真实生活。</p>\n<p>大多数人并不会突然从上班切换到退休。更多时候，是在某个阶段开始拥有更多选择。可以拒绝不合理的工作安排，可以选择收入低一点但更喜欢的方向，可以给家庭和健康留出更多空间，也可以把个人项目慢慢做起来。</p>\n<p>如果说 ChooseFIRE 想传达什么，我更希望它传达的是这种选择感。财务自由不只是一个终点数字，也是一套帮助人理解自己处境的方法。</p>\n<h2>从一个数字开始</h2>\n<p>如果对 FIRE 感兴趣，我觉得不必一上来就问“我什么时候可以退休”。</p>\n<p>更好的起点可能是三个数字：</p>\n<ul>\n<li>现在每年大概花多少钱？</li>\n<li>按这个支出水平，需要多少资产才能覆盖？</li>\n<li>以现在的储蓄速度，离这个目标还有多远？</li>\n</ul>\n<p>这三个问题已经足够把很多模糊感变清楚。</p>\n<p>接下来再去调整其他变量：支出高一点会怎样，低一点会怎样；收益率保守一点会怎样；每月多储蓄一点会怎样。算几轮之后，FIRE 就不再只是网上一个诱人的概念，而会变成和自己生活有关的一组取舍。</p>\n<p>这也是我做 <a href=\"https://choosefire.com/\">ChooseFIRE</a> 的原因。</p>\n<p>它不能替代投资判断，也不能替代人生选择。但如果它能帮人第一次认真算清楚自己的收入、支出、资产和时间之间的关系，这个工具就有价值。</p>\n<p>财务自由不该只靠感觉。至少，先把数字摆出来看看。</p>\n","date_published":"2026-05-08T00:00:00.000Z","tags":["FIRE","财务自由","提前退休","个人项目"],"language":"zh"},{"id":"https://www.lihuanyu.com/posts/deepseek-api-sillytavern-no-gpu/","url":"https://www.lihuanyu.com/posts/deepseek-api-sillytavern-no-gpu/","title":"DeepSeek API 接入 SillyTavern：不用本地显卡的小酒馆方案","summary":"在没有高性能显卡的情况下，通过 DeepSeek 官方 API 和 SillyTavern 的 Chat Completion / OpenAI-compatible 配置，把小酒馆连接到云端 DeepSeek 模型。","content_html":"<p>把 DeepSeek 接入 SillyTavern 小酒馆，有两条常见路线。</p>\n<p>第一条是本地部署：用 Ollama、KoboldCPP 或 LM Studio 在自己的电脑上跑模型，再让 SillyTavern 连接本机服务。这个方案适合有显卡、想离线使用、愿意折腾模型的人。</p>\n<p>第二条是 API：SillyTavern 仍然跑在本机，但模型推理交给 DeepSeek 官方 API。这个方案不需要本地显卡，也不用下载几十 GB 的模型文件。只要能稳定访问 API，普通电脑也可以使用。</p>\n<p>这篇记录第二种方案。需要本地 Ollama 方案的话，可以看 <a href=\"/posts/2025/%E6%9C%AC%E5%9C%B0%E9%83%A8%E7%BD%B2deepseek%E4%B8%8ESillyTavern/\">DeepSeek R1 接入 SillyTavern 小酒馆：Ollama 本地部署教程</a>。</p>\n<p>如果还没决定走本地还是 API，可以先看 <a href=\"/pages/sillytavern-deepseek/\">DeepSeek 接入 SillyTavern 指南：本地 Ollama 与 API 怎么选</a>。</p>\n<p><img src=\"/assets/posts/2026/deepseek-sillytavern-api/01-local-tavern-cloud-model.jpg\" alt=\"本机小酒馆界面通过安全通道连接云端模型服务\"></p>\n<h2>适合什么人</h2>\n<p>API 方案更适合这些情况：</p>\n<ul>\n<li>电脑没有独立显卡，或者显存不足。</li>\n<li>不想下载和管理本地模型文件。</li>\n<li>希望模型效果更接近官方网页。</li>\n<li>可以接受按 token 计费。</li>\n<li>主要在联网环境下使用小酒馆。</li>\n</ul>\n<p>它不适合追求完全离线的人。所有对话都会发送到模型服务商，角色卡、聊天内容和系统提示词都要按云端 API 的使用边界来理解。涉及隐私或敏感内容时，应先评估风险。</p>\n<h2>准备 DeepSeek API Key</h2>\n<p>先进入 DeepSeek 官方开放平台：</p>\n<p><a href=\"https://platform.deepseek.com/\">DeepSeek API Platform</a></p>\n<p>注册、登录、完成必要的账户设置后，创建 API Key。Key 通常只在创建时完整显示一次，保存时要放在密码管理器或本机安全位置，不要提交到 Git 仓库，也不要发到聊天记录里。</p>\n<p>模型和价格以官方文档为准：</p>\n<p><a href=\"https://api-docs.deepseek.com/quick_start/pricing\">DeepSeek Models &amp; Pricing</a></p>\n<p>截至 2026 年 7 月 31 日，DeepSeek 官方价格页列出的模型名是：</p>\n<ul>\n<li><code>deepseek-v4-flash</code></li>\n<li><code>deepseek-v4-pro</code></li>\n</ul>\n<p><code>deepseek-chat</code> 和 <code>deepseek-reasoner</code> 这两个旧模型名已经不在当前价格页的模型列表中。新配置直接使用 <code>deepseek-v4-flash</code> 或 <code>deepseek-v4-pro</code>，不要再照着旧截图填写。模型名可能继续变化，实际配置前仍应打开上面的官方页面确认一次。</p>\n<p>日常角色聊天可以先从 <code>deepseek-v4-flash</code> 开始。它通常更适合作为默认选择；如果对回复质量要求更高，再换成 <code>deepseek-v4-pro</code> 做对比。</p>\n<h2>安装并启动 SillyTavern</h2>\n<p>如果还没有安装 SillyTavern，Windows 上先安装：</p>\n<ul>\n<li><a href=\"https://git-scm.com/\">Git</a></li>\n<li><a href=\"https://nodejs.org/\">Node.js LTS</a></li>\n</ul>\n<p>然后在命令行执行：</p>\n<pre><code class=\"language-bash\">git clone https://github.com/SillyTavern/SillyTavern -b release\n</code></pre>\n<p>进入 <code>SillyTavern</code> 文件夹，双击 <code>Start.bat</code>。浏览器打开后，SillyTavern 本体就运行起来了。</p>\n<p>官方安装文档：<a href=\"https://docs.sillytavern.app/installation/windows/\">SillyTavern Windows Installation</a></p>\n<h2>配置 DeepSeek API</h2>\n<p>进入 SillyTavern 后，点击顶部的插头图标，打开 API 连接设置。不同版本的 UI 名称可能会有细微变化，但核心思路是：选择 Chat Completion，然后把 DeepSeek 当作云端 API 或 OpenAI-compatible API 接入。</p>\n<p><img src=\"/assets/posts/2026/deepseek-sillytavern-api/02-api-configuration-path.jpg\" alt=\"API Key、网络入口、模型选择和测试回复组成的配置路径\"></p>\n<h3>方式一：使用内置 DeepSeek 入口</h3>\n<p>如果当前 SillyTavern 版本的 Chat Completion Source 里已经有 DeepSeek，可以优先用这个入口：</p>\n<ul>\n<li>API 类型：Chat Completion</li>\n<li>Source / Provider：DeepSeek</li>\n<li>API Key：填入 DeepSeek 控制台生成的 Key</li>\n<li>Model：优先选择 <code>deepseek-v4-flash</code>，需要更强效果时选择 <code>deepseek-v4-pro</code></li>\n</ul>\n<p>保存后测试连接。能返回模型信息或测试回复，就说明配置成功。</p>\n<p>如果模型列表里仍然只有 <code>deepseek-chat</code>、<code>deepseek-reasoner</code> 这类旧名称，说明 SillyTavern 的内置列表可能还没跟上 DeepSeek 文档变化。此时可以改用下面的 OpenAI-compatible 方式，手动填写模型名。</p>\n<h3>方式二：使用 OpenAI-compatible 配置</h3>\n<p>DeepSeek API 兼容 OpenAI API 格式，因此也可以走 SillyTavern 的 Custom / OpenAI-compatible 配置。</p>\n<p>常用配置如下：</p>\n<pre><code class=\"language-text\">API 类型：Chat Completion\nSource / Provider：Custom 或 OpenAI-compatible\nAPI Key：sk-...\nBase URL：https://api.deepseek.com\nModel：deepseek-v4-flash\n</code></pre>\n<p>如果当前 SillyTavern 版本要求 OpenAI 风格的 <code>/v1</code> 地址，可以把 Base URL 改成：</p>\n<pre><code class=\"language-text\">https://api.deepseek.com/v1\n</code></pre>\n<p>不要把地址填成 <code>https://api.deepseek.com/chat/completions</code>。SillyTavern 会自己拼接具体接口路径，配置里通常只需要填基础地址。</p>\n<h2>推荐参数</h2>\n<p>角色聊天最重要的不是单个参数绝对正确，而是先让连接稳定，再慢慢调体验。可以从比较保守的配置开始：</p>\n<ul>\n<li>Model：<code>deepseek-v4-flash</code></li>\n<li>Temperature：<code>0.8</code> 到 <code>1.0</code></li>\n<li>Top P：<code>0.9</code></li>\n<li>Max response length：先设中等长度，确认回复速度和费用后再加大</li>\n<li>Streaming：开启，方便边生成边看</li>\n</ul>\n<p>如果角色说话过于发散，降低 Temperature；如果回复太短，增加最大回复长度；如果上下文费用增长太快，减少保留消息数量或缩短角色卡描述。</p>\n<h2>常见问题</h2>\n<h3>为什么 API 方案不需要显卡？</h3>\n<p>模型运行在 DeepSeek 的服务器上，本机只负责运行 SillyTavern 界面、发送请求和展示回复。因此普通笔记本也能使用，瓶颈主要变成网络、API 可用性和费用。</p>\n<h3>DeepSeek API 和本地 Ollama 版本有什么区别？</h3>\n<p>Ollama 运行的是本地模型，优点是可控、可以离线、没有按 token 计费；缺点是硬件要求高，模型越大越吃显存和内存。</p>\n<p>DeepSeek API 使用云端模型，优点是不用本地显卡、效果通常更稳定；缺点是联网依赖、按量计费，并且对话会发送到服务商。</p>\n<h3>填了 API Key 还是连不上怎么办？</h3>\n<p>优先排查四件事：</p>\n<ul>\n<li>API Key 是否复制完整，前后有没有多余空格。</li>\n<li>Base URL 是否只填基础地址，而不是完整接口路径。</li>\n<li>模型名是否是 DeepSeek 当前官方文档里的可用模型。</li>\n<li>网络是否能访问 DeepSeek API。</li>\n</ul>\n<p>如果内置 DeepSeek 入口失败，可以换 OpenAI-compatible 配置；如果 <code>https://api.deepseek.com</code> 不通，再尝试 <code>https://api.deepseek.com/v1</code>。</p>\n<h3>费用怎么控制？</h3>\n<p>角色聊天很容易因为上下文不断变长而增加 token 消耗。可以从这几件事控制：</p>\n<p><img src=\"/assets/posts/2026/deepseek-sillytavern-api/03-cost-privacy-network.jpg\" alt=\"网络可用性、费用仪表和隐私边界之间的 API 使用取舍\"></p>\n<ul>\n<li>先用 <code>deepseek-v4-flash</code>。</li>\n<li>不要一次保留过长聊天历史。</li>\n<li>控制角色卡、世界书和系统提示词长度。</li>\n<li>先短时间测试，再长期使用。</li>\n<li>定期查看 DeepSeek 控制台里的用量。</li>\n</ul>\n<h3>还应该保留本地部署方案吗？</h3>\n<p>可以保留。本地方案更像技术玩具和隐私偏好的选择，API 方案更像稳定使用的选择。实际体验后，我更倾向于把 API 作为日常小酒馆方案，把 Ollama 本地模型作为测试、离线和模型对比方案。</p>\n<h2>小结</h2>\n<p>DeepSeek API 接入 SillyTavern 的核心只有三件事：拿到 API Key，选择 Chat Completion / OpenAI-compatible，把 Base URL 和模型名填对。</p>\n<p>如果只是想在小酒馆里稳定使用 DeepSeek，不必先买显卡，也不必下载本地模型。API 方案的门槛更低，后续真正需要离线或本地可控时，再回到 Ollama、KoboldCPP 或 LM Studio 也不迟。</p>\n<h2>参考资料</h2>\n<ul>\n<li><a href=\"https://api-docs.deepseek.com/\">DeepSeek API Documentation</a></li>\n<li><a href=\"https://api-docs.deepseek.com/quick_start/pricing\">DeepSeek Models &amp; Pricing</a></li>\n<li><a href=\"https://docs.sillytavern.app/installation/windows/\">SillyTavern Windows Installation</a></li>\n<li><a href=\"https://docs.sillytavern.app/usage/api-connections/\">SillyTavern API Connections</a></li>\n</ul>\n<p><a href=\"/en/posts/2026/deepseek-api-sillytavern-no-gpu/\">English version: SillyTavern DeepSeek API Setup: Step-by-Step, No GPU Required</a></p>\n","date_published":"2026-05-05T00:00:00.000Z","date_modified":"2026-07-31T00:00:00.000Z","tags":["AI","DeepSeek","SillyTavern","小酒馆","API","教程"],"language":"zh"},{"id":"https://www.lihuanyu.com/en/posts/2026/deepseek-api-sillytavern-no-gpu/","url":"https://www.lihuanyu.com/en/posts/2026/deepseek-api-sillytavern-no-gpu/","title":"SillyTavern DeepSeek API Setup: Step-by-Step, No GPU Required","summary":"Configure the DeepSeek API in SillyTavern step by step with the built-in provider or an OpenAI-compatible connection, without a local GPU.","content_html":"<p>There are two common routes for using DeepSeek in SillyTavern.</p>\n<p>The first route is local deployment: run a model on your own computer with Ollama, KoboldCPP, or LM Studio, then let SillyTavern connect to that local service. This is good for people who have a GPU, want offline use, or enjoy comparing local models.</p>\n<p>The second route is the API route: SillyTavern still runs on your machine, but model inference happens through the official DeepSeek API. This does not require a local GPU, and it avoids downloading tens of gigabytes of model files. As long as the network and API account are usable, an ordinary laptop can run the SillyTavern side.</p>\n<p>This article covers the second route. If you want the local Ollama setup instead, see <a href=\"/en/posts/2025/deepseek-r1-sillytavern-ollama-local-deployment/\">DeepSeek R1 with SillyTavern: Local Ollama Setup on Windows</a>.</p>\n<p>For a side-by-side decision before configuring either route, see <a href=\"/en/pages/sillytavern-deepseek/\">SillyTavern with DeepSeek: Local Ollama vs API Setup</a>.</p>\n<p><img src=\"/assets/posts/2026/deepseek-sillytavern-api/01-local-tavern-cloud-model.jpg\" alt=\"A local SillyTavern interface connected to a cloud model service through a secure path\"></p>\n<h2>Quick Configuration</h2>\n<div class=\"table-scroll\"><table>\n<thead>\n<tr>\n<th>Setting</th>\n<th>Starting value</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>API key</td>\n<td>Create one in the DeepSeek API platform</td>\n</tr>\n<tr>\n<td>SillyTavern mode</td>\n<td>Chat Completion</td>\n</tr>\n<tr>\n<td>Provider</td>\n<td>DeepSeek, or Custom/OpenAI-compatible</td>\n</tr>\n<tr>\n<td>Base URL</td>\n<td><code>https://api.deepseek.com</code></td>\n</tr>\n<tr>\n<td>Alternative base URL</td>\n<td><code>https://api.deepseek.com/v1</code> when the current SillyTavern version requires it</td>\n</tr>\n<tr>\n<td>Model</td>\n<td>Use a current model listed in the official DeepSeek API documentation</td>\n</tr>\n</tbody>\n</table>\n</div><p>Do not paste the full <code>/chat/completions</code> endpoint into the Base URL field. SillyTavern builds the final request path. Provider labels and model names can change between versions, so use the detailed steps below together with the current official model list.</p>\n<h2>Who This Is For</h2>\n<p>The API route fits these situations:</p>\n<ul>\n<li>The computer has no dedicated GPU, or the VRAM is not enough.</li>\n<li>You do not want to download and manage local model files.</li>\n<li>You want model behavior closer to the official online DeepSeek experience.</li>\n<li>Token-based billing is acceptable.</li>\n<li>You mainly use SillyTavern while connected to the internet.</li>\n</ul>\n<p>It is not for fully offline use. Prompts, role cards, chat content, and system messages are sent to the model provider. If a conversation contains private or sensitive information, treat the cloud API boundary as part of the risk model.</p>\n<h2>Prepare a DeepSeek API Key</h2>\n<p>Go to the official DeepSeek API platform first:</p>\n<p><a href=\"https://platform.deepseek.com/\">DeepSeek API Platform</a></p>\n<p>After registration, login, and any required account setup, create an API key. The full key is usually shown only once. Save it in a password manager or another safe local place. Do not commit it to Git and do not paste it into chat logs.</p>\n<p>Model names and pricing should be checked on the official documentation:</p>\n<p><a href=\"https://api-docs.deepseek.com/quick_start/pricing\">DeepSeek Models &amp; Pricing</a></p>\n<p>As of July 31, 2026, DeepSeek’s official Models &amp; Pricing page lists:</p>\n<ul>\n<li><code>deepseek-v4-flash</code></li>\n<li><code>deepseek-v4-pro</code></li>\n</ul>\n<p>The older model names <code>deepseek-chat</code> and <code>deepseek-reasoner</code> no longer appear in the current pricing-page model list. Use <code>deepseek-v4-flash</code> or <code>deepseek-v4-pro</code> for a new configuration instead of copying an older screenshot. Model names can change again, so confirm the current list on the official page before configuring SillyTavern.</p>\n<p>For everyday role chat, I would start with <code>deepseek-v4-flash</code>. It is a better default for cost and speed. If response quality matters more, compare it with <code>deepseek-v4-pro</code>.</p>\n<h2>Install and Start SillyTavern</h2>\n<p>If SillyTavern is not installed yet, install these on Windows:</p>\n<ul>\n<li><a href=\"https://git-scm.com/\">Git</a></li>\n<li><a href=\"https://nodejs.org/\">Node.js LTS</a></li>\n</ul>\n<p>Then run:</p>\n<pre><code class=\"language-bash\">git clone https://github.com/SillyTavern/SillyTavern -b release\n</code></pre>\n<p>Enter the <code>SillyTavern</code> folder and double-click <code>Start.bat</code>. Once the browser opens, the SillyTavern application itself is running.</p>\n<p>Official installation guide: <a href=\"https://docs.sillytavern.app/installation/windows/\">SillyTavern Windows Installation</a></p>\n<h2>Configure the DeepSeek API</h2>\n<p>In SillyTavern, click the plug icon at the top to open API connection settings. The exact UI labels may change between versions, but the idea is stable: choose Chat Completion, then connect DeepSeek either through the built-in DeepSeek provider or through an OpenAI-compatible custom provider.</p>\n<p><img src=\"/assets/posts/2026/deepseek-sillytavern-api/02-api-configuration-path.jpg\" alt=\"The configuration path made of API key, network endpoint, model choice, and test response\"></p>\n<h3>Option 1: Use the Built-In DeepSeek Provider</h3>\n<p>If your current SillyTavern version already includes DeepSeek as a Chat Completion source, start there:</p>\n<ul>\n<li>API type: Chat Completion</li>\n<li>Source / Provider: DeepSeek</li>\n<li>API Key: the key created in the DeepSeek platform</li>\n<li>Model: start with <code>deepseek-v4-flash</code>, or use <code>deepseek-v4-pro</code> when quality matters more</li>\n</ul>\n<p>Save the settings and test the connection. If SillyTavern returns model information or a test reply, the connection is working.</p>\n<p>If the model list still only shows older names such as <code>deepseek-chat</code> and <code>deepseek-reasoner</code>, SillyTavern’s built-in list may not have caught up with DeepSeek’s documentation. In that case, use the OpenAI-compatible route and type the model name manually.</p>\n<h3>Option 2: Use an OpenAI-Compatible Configuration</h3>\n<p>DeepSeek’s API is compatible with the OpenAI API shape, so SillyTavern can also connect through a Custom or OpenAI-compatible provider.</p>\n<p>A common configuration looks like this:</p>\n<pre><code class=\"language-text\">API type: Chat Completion\nSource / Provider: Custom or OpenAI-compatible\nAPI Key: sk-...\nBase URL: https://api.deepseek.com\nModel: deepseek-v4-flash\n</code></pre>\n<p>If the current SillyTavern version expects an OpenAI-style <code>/v1</code> base URL, use:</p>\n<pre><code class=\"language-text\">https://api.deepseek.com/v1\n</code></pre>\n<p>Do not set the base URL to <code>https://api.deepseek.com/chat/completions</code>. SillyTavern builds the final endpoint path itself. In the configuration field, it usually needs only the base address.</p>\n<h2>Suggested Starting Parameters</h2>\n<p>For role chat, there is no single perfect parameter set. The first goal is to make the connection stable, then tune the behavior.</p>\n<p>A conservative starting point:</p>\n<ul>\n<li>Model: <code>deepseek-v4-flash</code></li>\n<li>Temperature: <code>0.8</code> to <code>1.0</code></li>\n<li>Top P: <code>0.9</code></li>\n<li>Max response length: start medium, then increase after checking speed and cost</li>\n<li>Streaming: enabled, so responses appear while being generated</li>\n</ul>\n<p>If the character becomes too unfocused, lower Temperature. If replies are too short, increase max response length. If context cost grows too quickly, keep fewer chat messages or shorten the character card.</p>\n<h2>Common Questions</h2>\n<h3>Why does the API route not need a GPU?</h3>\n<p>The model runs on DeepSeek’s servers. Your computer only runs the SillyTavern interface, sends requests, and displays responses. The bottlenecks become network access, API availability, and cost, not local GPU power.</p>\n<h3>What is the difference between DeepSeek API and local Ollama?</h3>\n<p>Ollama runs a local model. It gives you more control, can work offline, and does not bill by token. The tradeoff is hardware pressure: larger models need more VRAM and RAM.</p>\n<p>DeepSeek API runs a cloud model. It needs no local GPU and is usually more stable. The tradeoffs are network dependency, usage-based billing, and the fact that conversation data is sent to the provider.</p>\n<h3>What if I entered the API key but it still cannot connect?</h3>\n<p>Check four things first:</p>\n<ul>\n<li>Whether the API key was copied completely, without extra spaces.</li>\n<li>Whether the Base URL is only the base address, not the full endpoint path.</li>\n<li>Whether the model name is currently available in DeepSeek’s official documentation.</li>\n<li>Whether the network can reach the DeepSeek API.</li>\n</ul>\n<p>If the built-in DeepSeek provider fails, try the OpenAI-compatible configuration. If <code>https://api.deepseek.com</code> does not work in your SillyTavern version, try <code>https://api.deepseek.com/v1</code>.</p>\n<h3>How can I control cost?</h3>\n<p>Role chat can consume more tokens over time because the conversation context keeps growing.</p>\n<p><img src=\"/assets/posts/2026/deepseek-sillytavern-api/03-cost-privacy-network.jpg\" alt=\"API usage tradeoffs between network availability, cost, and privacy boundaries\"></p>\n<p>Start with these habits:</p>\n<ul>\n<li>Use <code>deepseek-v4-flash</code> first.</li>\n<li>Avoid keeping an unnecessarily long chat history.</li>\n<li>Keep role cards, world info, and system prompts reasonably short.</li>\n<li>Test briefly before long sessions.</li>\n<li>Check usage in the DeepSeek console regularly.</li>\n</ul>\n<h3>Should I still keep a local deployment?</h3>\n<p>Yes, if it fits the way you use SillyTavern.</p>\n<p>The local route is more like a technical playground and a privacy preference. The API route is more like a stable daily-use option. After trying both, I prefer using the API route for regular SillyTavern sessions, and keeping Ollama for testing, offline use, and model comparison.</p>\n<h2>Summary</h2>\n<p>Connecting DeepSeek API to SillyTavern comes down to three things: get an API key, choose Chat Completion or OpenAI-compatible mode, and set the Base URL plus model name correctly.</p>\n<p>If the goal is simply to use DeepSeek in SillyTavern without buying a GPU or downloading local model files, the API route has a much lower entry cost. If offline control becomes important later, it is still easy to return to Ollama, KoboldCPP, or LM Studio.</p>\n<h2>References</h2>\n<ul>\n<li><a href=\"https://api-docs.deepseek.com/\">DeepSeek API Documentation</a></li>\n<li><a href=\"https://api-docs.deepseek.com/quick_start/pricing\">DeepSeek Models &amp; Pricing</a></li>\n<li><a href=\"https://docs.sillytavern.app/installation/windows/\">SillyTavern Windows Installation</a></li>\n<li><a href=\"https://docs.sillytavern.app/usage/api-connections/\">SillyTavern API Connections</a></li>\n</ul>\n<p><a href=\"/posts/deepseek-api-sillytavern-no-gpu/\">Chinese version of this article</a></p>\n","date_published":"2026-05-05T00:00:00.000Z","date_modified":"2026-07-31T00:00:00.000Z","tags":["AI","DeepSeek","SillyTavern","API","Tutorial"],"language":"en"},{"id":"https://www.lihuanyu.com/posts/2026/%E5%A2%A8%E5%B1%BF-InkIsle-%E6%88%91%E4%B8%BA%E4%BB%80%E4%B9%88%E5%8F%88%E5%86%99%E4%BA%86%E4%B8%80%E4%B8%AA%E5%8D%9A%E5%AE%A2%E7%B3%BB%E7%BB%9F/","url":"https://www.lihuanyu.com/posts/2026/%E5%A2%A8%E5%B1%BF-InkIsle-%E6%88%91%E4%B8%BA%E4%BB%80%E4%B9%88%E5%8F%88%E5%86%99%E4%BA%86%E4%B8%80%E4%B8%AA%E5%8D%9A%E5%AE%A2%E7%B3%BB%E7%BB%9F/","title":"墨屿 InkIsle：我为什么又写了一个博客系统","summary":"从 Hexo 个人博客迁移出发，介绍墨屿 InkIsle 的背景、设计取舍、实现方案、迁移结果、使用方式和未来开源计划。","content_html":"<p>2017 年，我把个人博客从 WordPress 换成了 Hexo。当时的判断很朴素：WordPress 对一个个人博客来说太重了，静态 HTML 更便宜、更安全，也更容易部署。</p>\n<p>几年之后，我又把 Hexo 换掉了。原因不是 Hexo 不好，而是我的需求变了。</p>\n<p>我希望博客仍然是 Markdown 驱动，仍然能静态输出，但写作、构建、主题、多语言、搜索、RSS、JSON Feed、<code>llms.txt</code> 这些东西应该作为一个整体被设计。尤其是 AI 时代，公开内容不只是给人读，也会被搜索引擎、RSS 阅读器和 AI agent 消费。于是我写了一个新的博客系统：墨屿，英文名 InkIsle。</p>\n<p><img src=\"/assets/posts/2026/inkisle-blog-system/01-markdown-island.jpg\" alt=\"Markdown 内容岛屿连接搜索、RSS、AI 和站点结构等出口\"></p>\n<h2>为什么想要一个新的博客系统</h2>\n<p>我对旧博客系统的不满主要有三个。</p>\n<p>第一是性能和发布体验。Hexo 是成熟方案，但在我的博客内容逐渐变多、迁移历史文章和多语言内容之后，构建和本地调试体验不够理想。我希望从提交到线上可见尽量控制在分钟级，平时本地写作也不要有明显等待。</p>\n<p>第二是主题和内容耦合。传统博客系统经常把主题、页面结构、插件和内容规则缠在一起。文章本身应该只是文章，最好放在清楚的 <code>content/</code> 目录里；主题应该负责视觉和布局；RSS、搜索、站点地图、AI 输出这些功能则应该属于系统能力，而不是某个主题顺手实现的东西。</p>\n<p>第三是 AI 友好。过去博客主要面向浏览器里的读者，现在公开内容也需要更容易被机器理解。<code>llms.txt</code>、结构化 JSON、搜索索引、清晰的多语言 URL、稳定的 Markdown 内容结构，都会变得越来越重要。</p>\n<p>所以我的目标不是“再换一个主题”，而是重新整理一套更适合长期使用的内容发布方式。</p>\n<h2>为什么自己写</h2>\n<p>这个问题其实很值得先问：博客系统已经很多了，为什么还要自己写？</p>\n<p>一开始我也看过现有方案。直接用 Astro starter、VitePress、Nextra、Nuxt Content、Eleventy 之类，都能做出不错的内容站。但我想要的不是一个单站点模板，而是一层更产品化的封装：</p>\n<ul>\n<li>默认只暴露 Markdown、配置和静态资源，不要求用户理解完整 Astro 工程结构。</li>\n<li>主题和 renderer 分开，个人博客和商业内容站可以用同一套内容模型。</li>\n<li>内建 RSS、sitemap、静态搜索、JSON Feed、posts JSON、<code>llms.txt</code>。</li>\n<li>支持主语言内容和翻译内容分目录组织。</li>\n<li>能作为 npm CLI 使用，而不是每个项目复制一份框架代码。</li>\n<li>未来可以复用到公司网站、商业博客、文档站和其他内容型产品。</li>\n</ul>\n<p>如果只是给自己的博客换个样式，直接改 Hexo 或换 Astro starter 就够了。但我想验证的是一套更清楚的发布产品层：Markdown 是内容源，Astro 是渲染底座，InkIsle 负责把这些能力包装成简单的工作流。</p>\n<p>这也是我选择自己写的原因。不是因为市面上没有成熟方案，而是因为我想要的边界和体验足够具体，自己实现一个最小系统反而更容易把它打磨成自己长期会用的东西。</p>\n<h2>为什么用 Astro</h2>\n<p>InkIsle 没有从零实现静态站点生成器，底层选择了 Astro。</p>\n<p>原因也很直接。</p>\n<p>Astro 默认适合静态输出，可以把 Markdown 预渲染成 HTML；它的构建体系基于 Vite，本地开发和构建性能都很好；Markdown、MDX、静态路由、动态路由、RSS、sitemap、adapter 这些能力都有成熟基础。后续如果真的需要 SSR 或 Edge rendering，也有 adapter 的扩展路径。</p>\n<p>我一开始也考虑过偏 Next.js 方向的方案，比如 Vinext 这类更开放的全栈框架探索。但博客系统的第一目标不是运行时应用，而是静态优先、预渲染优先、内容优先。评论、登录态、动态预览这些都可以晚点做，文章页本身应该尽量只是静态 HTML。</p>\n<p>所以最终的取舍是：不重写底层构建器，也不把博客做成复杂应用。Astro 做底座，InkIsle 做产品层。</p>\n<h2>InkIsle 的设计</h2>\n<p>InkIsle 的核心结构是两个 starter。</p>\n<p><img src=\"/assets/posts/2026/inkisle-blog-system/02-content-renderer-separation.jpg\" alt=\"内容层、渲染层、主题层和系统能力被分层组织\"></p>\n<p>默认的 <code>content-only</code> starter 面向普通使用者，项目里只有这些东西：</p>\n<pre><code class=\"language-text\">content/\npublic/\ninkisle.config.mjs\npackage.json\n</code></pre>\n<p>这意味着一个博客项目不需要看到 <code>astro.config.mjs</code>、<code>src/pages/</code>、布局组件和 renderer 实现细节。日常写作只需要维护 Markdown 内容和配置。</p>\n<p>另一个 <code>default</code> starter 是完整 Astro renderer，放在 InkIsle 包里。它负责读取内容、生成页面、套用主题、输出 feed、搜索索引和静态文件。</p>\n<p>内容结构大致是这样：</p>\n<pre><code class=\"language-text\">content/posts/my-post.md\ncontent/pages/about.md\ncontent/en/posts/my-post.md\ncontent/en/pages/about.md\n</code></pre>\n<p>主语言内容直接放在 <code>content/posts/</code> 和 <code>content/pages/</code>，翻译内容放在 <code>content/{lang}/</code> 下。默认情况下，主语言发布到根路径，英文等翻译内容使用语言前缀：</p>\n<pre><code class=\"language-text\">/posts/my-post/\n/en/posts/my-post/\n</code></pre>\n<p>这个选择和最初设想有一点调整。最初我想过构建后主语言也带 <code>/zh/</code> 前缀，后来迁移真实博客时发现，个人博客的主语言路径不带前缀更自然，也更利于保留旧链接和读者习惯。因此 InkIsle 默认主语言无前缀，同时保留 <code>/zh/...</code> 兼容重定向。</p>\n<h2>已经完成的能力</h2>\n<p>目前 InkIsle 已经完成了一个可以真实使用的版本，并且发布到了 npm。</p>\n<p>现在已经有：</p>\n<ul>\n<li><code>inkisle init</code> 创建默认内容站。</li>\n<li><code>inkisle init --full</code> 创建完整 Astro 项目。</li>\n<li><code>inkisle new post</code> 和 <code>inkisle new page</code> 创建内容。</li>\n<li><code>inkisle dev</code>、<code>inkisle build</code> 等命令跑本地开发、构建和本地预览。</li>\n<li><code>inkisle check links</code> 检查构建产物里的站内链接。</li>\n<li>个人博客主题 <code>personal</code>。</li>\n<li>商业内容站主题 <code>business-blog</code>。</li>\n<li>文章列表、分页、标签页、分类页、自定义页面、搜索页、404。</li>\n<li>RSS、JSON Feed、<code>/api/posts.json</code>、<code>/search-index.json</code>、<code>/llms.txt</code>、sitemap、robots.txt。</li>\n<li>多语言内容路径。</li>\n<li>默认语言无前缀，默认语言前缀路径兼容重定向。</li>\n<li>PWA manifest、service worker、备案号、百度统计、搜索验证文件、Cloudflare Pages <code>_redirects</code> 等站点级配置。</li>\n</ul>\n<p>有些能力还只是配置形态或规划方向，比如 raw Markdown 输出和单篇文章 JSON 输出。它们在最初设计里很重要，但不是迁移个人博客的第一优先级，所以我暂时没有强行把所有想法一次做完。</p>\n<h2>成品效果</h2>\n<p>我的个人博客现在已经迁移到 InkIsle。</p>\n<p><img src=\"/assets/posts/2026/inkisle-blog-system/03-migration-build-pipeline.jpg\" alt=\"旧博客内容通过迁移桥梁进入新的构建流水线和灯塔站点\"></p>\n<p>旧博客内容被整理成 <code>content/posts/</code>，英文翻译文章放在 <code>content/en/posts/</code>。站点配置集中在 <code>inkisle.config.mjs</code>，构建命令也变得很直接：</p>\n<pre><code class=\"language-bash\">pnpm run build\npnpm run check:links\npnpm run deploy\n</code></pre>\n<p>迁移后的实际构建结果比我预期更好。当前博客大约生成 400 多个 HTML 页面，构建耗时在几秒级；站内链接检查会扫描数千个链接，用来避免迁移历史文章时留下坏链接。</p>\n<p>更重要的是，生成结果不只是网页：</p>\n<ul>\n<li><code>/rss.xml</code> 给 RSS 阅读器。</li>\n<li><code>/feed.json</code> 给 JSON Feed 读者。</li>\n<li><code>/api/posts.json</code> 给结构化消费。</li>\n<li><code>/search-index.json</code> 给站内搜索。</li>\n<li><code>/llms.txt</code> 给 AI agent 一个入口。</li>\n<li><code>/sitemap-index.xml</code> 给搜索引擎。</li>\n</ul>\n<p>这正是我想要的新博客系统：不是单纯把 Markdown 变成网页，而是把公开内容整理成多个稳定、机器可读、长期可维护的出口。</p>\n<p><img src=\"/assets/posts/2026/inkisle-blog-system/04-ai-readable-archive.jpg\" alt=\"内容档案被整理成可供读者、RSS、搜索和 AI agent 访问的多出口系统\"></p>\n<h2>Logo 概念稿</h2>\n<p>后来我也顺手给 InkIsle 跑了一版 Logo 概念稿。</p>\n<p>我给模型的关键词是“墨迹、岛屿、Markdown 页面、静态发布、人与 AI 共同阅读”。这张图里，一滴墨形成了一座小岛，岛上立着一张像 Markdown 文档的白色页面，远处有一条细线和一个红点，像海岸线、发布路径，也像最后一次确认。</p>\n<p><img src=\"/assets/posts/2026/inkisle-blog-system/05-inkisle-concept-logo.jpg\" alt=\"墨屿 InkIsle 观念叙事型极简 Logo 概念稿\"></p>\n<p>它还不能直接当正式 Logo 用。图片模型给出的更像方向稿，真要放进系统里，还需要把图形抽出来，做成可缩放的 SVG，再检查 favicon、深浅色背景、小尺寸和印刷场景。至少方向上我挺喜欢：它没有把博客系统画成电脑、云朵或代码括号，而是回到了“墨”和“岛”这两个字本身。</p>\n<h2>如果你也想用</h2>\n<p>目前 npm 包已经发布，可以直接试用：</p>\n<pre><code class=\"language-bash\">npm exec inkisle -- init my-blog\ncd my-blog\nnpm install\nnpm run dev\n</code></pre>\n<p>创建文章：</p>\n<pre><code class=\"language-bash\">npm exec inkisle -- new post &quot;我的第一篇文章&quot; --published\n</code></pre>\n<p>构建：</p>\n<pre><code class=\"language-bash\">npm run build\n</code></pre>\n<p>如果你想看到完整 Astro 工程，而不是默认的轻量内容站，可以用：</p>\n<pre><code class=\"language-bash\">npm exec inkisle -- init my-full-blog --full\n</code></pre>\n<p>不过我更推荐从默认模式开始。InkIsle 的设计目标就是让多数使用者只关心内容、配置和资源，不必一上来面对完整前端工程。</p>\n<h2>未来开源计划</h2>\n<p>InkIsle 现在还处在早期阶段。代码目前主要服务我的个人博客迁移和真实验证，后续我会在这些方面继续补：</p>\n<ul>\n<li>完善 README 和使用文档。</li>\n<li>补 raw Markdown 和单篇 JSON 输出。</li>\n<li>梳理主题 API，决定什么时候支持本地主题和 npm 主题。</li>\n<li>增加更清楚的内容质量检查。</li>\n<li>改进多语言翻译工作流，至少给 AI 翻译提供明确的路径约定。</li>\n<li>为 Cloudflare Pages、GitHub Pages 和自托管部署补示例。</li>\n<li>等 API 和主题边界稳定后公开仓库。</li>\n</ul>\n<p>我不想太早把它包装成一个“通用框架”。更合理的路径是先服务真实博客，把个人长期使用中遇到的问题都解决掉，再把稳定的部分开源出来。</p>\n<h2>写博客系统这件事</h2>\n<p>写博客系统听起来像一种重复造轮子。很多时候也确实是。</p>\n<p>但个人博客有一个很特殊的地方：它既是工具，也是自己的写作空间。工具的边界会反过来影响写作习惯、内容整理方式、发布频率和长期维护成本。</p>\n<p>2017 年从 WordPress 换到 Hexo，是为了从动态博客走向静态博客。2026 年从 Hexo 换到 InkIsle，是为了从“能生成网页”走向“更适合人和 AI 一起消费的 Markdown 发布系统”。</p>\n<p>这个目标听起来不大，但足够具体。对我来说，一个会长期使用、能承载自己内容资产、还能慢慢变成开源产品的博客系统，值得认真写一次。</p>\n","date_published":"2026-05-05T00:00:00.000Z","tags":["InkIsle","Astro","Markdown","独立博客","AI"],"language":"zh"},{"id":"https://www.lihuanyu.com/posts/2026/AI%E6%97%B6%E4%BB%A3%E5%89%8D%E5%90%8E%E7%AB%AF%E5%88%86%E7%A6%BB%E4%B8%8D%E8%AF%A5%E5%86%8D%E6%98%AF%E9%BB%98%E8%AE%A4%E9%80%89%E9%A1%B9/","url":"https://www.lihuanyu.com/posts/2026/AI%E6%97%B6%E4%BB%A3%E5%89%8D%E5%90%8E%E7%AB%AF%E5%88%86%E7%A6%BB%E4%B8%8D%E8%AF%A5%E5%86%8D%E6%98%AF%E9%BB%98%E8%AE%A4%E9%80%89%E9%A1%B9/","title":"AI时代，前后端分离不该再是默认选项","summary":"从 AI 对上下文完整性的依赖出发，重新讨论前后端分离、产品层全栈、同构框架和业务域团队在新时代的默认选择。","content_html":"<p>我过去对前后端分离的看法是：前后端应该分离，但开发前后端的人不应该分离。</p>\n<p>意思是，系统边界可以拆，接口可以清楚，前端和后端可以有不同的工程结构；但真正负责一个功能的人，最好理解从页面到数据、从交互到业务规则的完整链路。否则前端只知道调接口，后端只知道吐 JSON，最后很容易变成每个人都只对自己那一段负责，却没人真正对用户体验和业务结果负责。</p>\n<p>现在我的看法又往前走了一步：在 AI 时代，很多项目连前后端本身都不应该默认分离了。更准确地说，产品层的前后端不该再默认分离。</p>\n<p><a href=\"/en/posts/2026/frontend-backend-separation-should-not-be-default-ai-era/\">English version: In the AI Era, Frontend-Backend Separation Should No Longer Be the Default</a></p>\n<p><img src=\"/assets/posts/2026/frontend-backend-context/01-split-context-bridge.jpg\" alt=\"两个产品上下文被窄桥连接，象征前后端分离带来的上下文断裂\"></p>\n<h2>什么是产品层全栈</h2>\n<p>这里说的产品层全栈，不是说所有系统都要塞进一个巨大的单体里，也不是否认底层平台、核心系统、数据能力的价值。它指的是一个面向用户的业务功能，应该尽量在同一个上下文里完成：界面、交互、数据读取、权限、状态、提交、校验、业务规则、持久化，以及部署发布。</p>\n<p><img src=\"/assets/posts/2026/frontend-backend-context/02-product-layer-full-stack.jpg\" alt=\"界面、权限、数据和部署被组织进同一个产品层空间\"></p>\n<p>对于小中型项目，这基本就是整个业务本身。对于大型系统，它也可以是围绕某个业务域的完整产品闭环。至于支付、登录、搜索、推荐、风控、数据分析、数仓等能力，如果复杂到需要独立演进，可以作为平台能力或业务依赖存在。</p>\n<p>换句话说，用户如何下单、课程如何售卖、文章如何发布、任务如何流转，这些仍然是产品业务逻辑，应该尽量留在产品层的完整上下文里。底层能力可以拆出去，但一个产品功能不应该天然按“前端”和“后端”切成两个互相等待的半成品。</p>\n<h2>前后端分离曾经是合理的</h2>\n<p>过去前后端分离解决的是人的问题。</p>\n<p>前端和后端技术栈差异大，关注点不同，团队规模变大后需要协作边界，接口契约可以减少沟通成本，独立部署也能降低互相影响。在工具不够强、个人跨栈成本较高的时候，这些理由都成立。</p>\n<p>但很多团队后来把它变成了一种默认先进性：只要做 Web 应用，就先拆前端项目和后端项目；只要有页面数据，就先设计 REST 或 GraphQL 接口；只要有前端和后端岗位，就默认两拨人分别负责。久而久之，“前端不应该碰后端，后端不应该写页面”也变成了一种近乎本能的组织假设。</p>\n<p>这个假设在 AI 时代变得越来越可疑。</p>\n<h2>AI 需要完整上下文</h2>\n<p>AI 改变的不是某个框架细节，而是开发者处理上下文的能力。AI 写代码的质量高度依赖上下文完整性。</p>\n<p>一个功能的页面、数据结构、权限判断、提交逻辑、错误处理、缓存策略、测试用例，如果都在同一个项目、同一种类型系统和相近的文件结构里，AI 可以更容易理解因果关系，也更容易做出连贯修改。</p>\n<p><img src=\"/assets/posts/2026/frontend-backend-context/03-ai-complete-context.jpg\" alt=\"完整上下文被汇聚到发光的中心模型，周围连接页面、数据、测试和规则\"></p>\n<p>反过来，如果前端在一个仓库，后端在另一个仓库；前端只有接口文档，后端看不到页面真实用法；部署链路也分开；类型还要通过生成代码或文档同步，那么人和 AI 都需要不断在局部信息之间补全上下文。这个成本过去由人承担，现在也会直接影响 AI 的产出质量。</p>\n<p>这也是我重新理解同构全栈框架的原因。</p>\n<h2>同构框架的价值不是只在 SSR</h2>\n<p>Next.js 这类框架的价值，不只是 SSR 能改善首屏体验，也不只是把 API route 和页面放在一起。更重要的是，它让一个产品功能的上下文重新合并了。</p>\n<p>页面怎么展示、数据怎么读、提交动作怎么处理、权限在哪里判断、类型怎么流动、缓存怎么失效，这些事情可以在同一个工程模型里协同。对开发者来说，这是心智负担的降低；对 AI 来说，这是上下文质量的提升。</p>\n<p>我个人也更喜欢 Vinext 代表的方向：保留 Next.js 这种全栈同构开发体验，同时尝试把构建和运行时放到更开放的 Vite、Cloudflare Workers 等生态里。当然，Vinext 现在仍然偏实验性，框架本身也不是重点。重点是这个趋势：开发上下文正在重新合并，全栈框架会越来越围绕 AI 友好的工程形态演进。</p>\n<p>Django、Rails、Laravel 这类传统全栈框架当然也属于全栈路线。它们长期证明了“一个项目完成产品功能”并不是什么落后的做法。只是对于前端交互复杂、组件生态依赖较重的现代 Web 应用，Next.js、Vinext 这类同构方案更贴近当前前端工程的工作方式。</p>\n<h2>接口契约没有消失</h2>\n<p>有人会说，前后端分离的价值在于接口契约。</p>\n<p>契约仍然重要，但它不一定需要表现为一个只服务本项目、却伪装成公共服务的 HTTP API。外部系统、移动端、多端复用、第三方集成，当然需要稳定 API。但 Web UI 自己消费的数据接口，不必天然被设计成公共契约。</p>\n<p>产品层全栈之后，契约不是消失了，而是内化了。它可以是 TypeScript 类型、schema、server action、组件 props、数据库模型、单元测试、集成测试和端到端测试。相比一份前后端隔着仓库维护的接口文档，这些契约更贴近代码真实运行的位置，也更容易被 AI 和工具一起理解。</p>\n<p>数据库也是类似的问题。小项目里，产品层直接读写数据库很正常；大项目里，可以通过领域服务、存储服务或平台能力访问数据。但不应该为了“前后端分离”而强行包装一层只服务页面的 API。</p>\n<h2>BFF 的位置也变了</h2>\n<p>BFF 也是类似的问题。</p>\n<p>很多 BFF 实践，本质上是在前后端分离之后补出来的胶水层：转发接口、裁剪字段、拼装数据、做一点鉴权和格式转换。它的存在反过来说明了一件事：纯 API 并不能很好地服务页面体验。既然如此，很多低价值 BFF 不如直接收回到产品层全栈里。</p>\n<p>当然，BFF 不是永远没有价值。如果它承担的是复杂聚合、安全治理、流量控制、缓存、灰度、降级，或者多个产品共享的体验编排，那它可以成为独立系统。但如果它只是为了让前端不要碰后端而存在，那它很可能只是组织边界制造出的额外复杂度。</p>\n<h2>团队也应该按业务域组织</h2>\n<p>更合理的组织方式，也不应该继续按前端和后端切人，而应该按业务域组织小型全栈团队。</p>\n<p>一个业务域里的工程师可以有专长，有人更熟悉交互，有人更熟悉数据，有人更熟悉基础设施；但团队整体应该拥有从产品需求到上线运行的完整闭环。长期按技术栈切人，会让工程师只理解链路的一段，最终没人真正理解用户需求如何落地。</p>\n<p><img src=\"/assets/posts/2026/frontend-backend-context/04-business-domain-team.jpg\" alt=\"多个业务域围绕共享中枢组织成小型全栈团队\"></p>\n<p>这并不意味着所有项目都不该拆。多端复用同一套 API、开放平台、复杂核心业务系统、强安全合规、极致性能、超大团队协作、顶级大厂面向 C 端的应用，这些场景仍然可能需要清晰的前后端边界。尤其是多端场景，同一套业务能力需要服务 Web、iOS、Android、小程序、第三方系统时，稳定 API 的价值会明显上升。</p>\n<p>但这些应该是拆分的理由，而不是默认起点。</p>\n<h2>一点个人实践感受</h2>\n<p>我的实践感受也很直接。以前做个人项目时，我常用 NestJS 加一个 Web 页面。这个组合不是不能用，但接口定义、类型同步、联调、部署、文件跳转、上下文切换，都会不断出现。</p>\n<p>后来换到 Vinext 这类全栈同构方案，最明显的变化不是少写了几行代码，而是一个功能终于能在一个上下文里完成。对人是这样，对 AI 更是这样。</p>\n<p>这不是什么大规模项目里的严谨结论，更像是一个开发者在实际写代码时积累出来的取舍感受。但很多架构判断，最终也会回到这种朴素的问题：一个功能到底是在帮助人更快地理解业务，还是在制造更多需要同步的边界？</p>\n<h2>默认全栈，除非有理由拆开</h2>\n<p>所以现在我更倾向于一个新的默认判断：</p>\n<ul>\n<li>如果只有一个主要 Web 端，优先产品层全栈。</li>\n<li>如果页面数据强依赖 UI 形态，优先产品层全栈。</li>\n<li>如果 API 主要服务自己的页面，而不是外部消费者，优先产品层全栈。</li>\n<li>如果团队规模还没大到需要强边界，优先产品层全栈。</li>\n<li>如果业务复杂了，先按业务域和平台能力拆，而不是先按前端和后端拆。</li>\n</ul>\n<p>以后讨论架构时，不应该再先问“要不要前后端分离”，而应该先问：</p>\n<p>这个业务有没有足够强的理由，把产品上下文拆开？</p>\n<p>如果没有，前后端分离就不该再是默认选项。它不是先进性的象征，而是一种需要证明必要性的复杂度。</p>\n","date_published":"2026-05-04T00:00:00.000Z","tags":["AI","全栈","前后端分离","架构"],"language":"zh"},{"id":"https://www.lihuanyu.com/en/posts/2026/frontend-backend-separation-should-not-be-default-ai-era/","url":"https://www.lihuanyu.com/en/posts/2026/frontend-backend-separation-should-not-be-default-ai-era/","title":"In the AI Era, Frontend-Backend Separation Should No Longer Be the Default","summary":"A reflection on why AI changes the cost-benefit balance of frontend-backend separation, and why product-layer full stack should become the default for many projects.","content_html":"<p>I used to think about frontend-backend separation this way: the frontend and backend could be separated, but the people building them should not be separated.</p>\n<p>In other words, system boundaries can exist. Interfaces can be clear. Frontend and backend code can have different structures. But the person responsible for a feature should understand the full path from page to data, from interaction to business rule. Otherwise, frontend engineers only call APIs, backend engineers only return JSON, and everyone ends up responsible for one slice of the chain while nobody is truly responsible for the user experience or the business result.</p>\n<p>My view has moved further now. In the AI era, many projects should no longer treat frontend-backend separation itself as the default. More precisely, the product layer should not default to frontend-backend separation.</p>\n<p><a href=\"/posts/2026/AI%E6%97%B6%E4%BB%A3%E5%89%8D%E5%90%8E%E7%AB%AF%E5%88%86%E7%A6%BB%E4%B8%8D%E8%AF%A5%E5%86%8D%E6%98%AF%E9%BB%98%E8%AE%A4%E9%80%89%E9%A1%B9/\">Chinese version of this article</a></p>\n<p><img src=\"/assets/posts/2026/frontend-backend-context/01-split-context-bridge.jpg\" alt=\"Two product contexts connected by a narrow bridge, representing the fragmented context created by frontend-backend separation\"></p>\n<h2>What Product-Layer Full Stack Means</h2>\n<p>Product-layer full stack does not mean putting every system into one giant monolith. It also does not deny the value of platform capabilities, core systems, or data infrastructure.</p>\n<p>It means that a user-facing product feature should, as much as possible, be completed in one context: interface, interaction, data access, permissions, state, submission, validation, business rules, persistence, and deployment.</p>\n<p><img src=\"/assets/posts/2026/frontend-backend-context/02-product-layer-full-stack.jpg\" alt=\"Interface, permissions, data, and deployment arranged inside one product-layer space\"></p>\n<p>For small and medium-sized projects, this is often the whole business. For larger systems, it can still be the complete product loop around one business domain. Capabilities such as payment, login, search, recommendation, risk control, analytics, and data warehouses can become platform capabilities or business dependencies when they are complex enough to evolve independently.</p>\n<p>Put differently, how users place orders, how courses are sold, how articles are published, or how tasks move through a workflow are still product business logic. They should usually stay inside the product layer’s complete context. Lower-level capabilities can be split out, but a product feature should not naturally be cut into two half-finished pieces called “frontend” and “backend”.</p>\n<h2>Separation Used to Make Sense</h2>\n<p>Frontend-backend separation used to solve a people problem.</p>\n<p>Frontend and backend stacks were different. Their concerns were different. As teams grew, they needed collaboration boundaries. Interface contracts reduced coordination cost. Independent deployment could reduce mutual impact. When tools were weaker and cross-stack work was expensive for individuals, those arguments made sense.</p>\n<p>But many teams later turned this into a default sign of engineering maturity. As soon as a Web app was built, there would be a frontend project and a backend project. As soon as a page needed data, a REST or GraphQL API would be designed. As soon as there were frontend and backend roles, two groups of people would own the two halves by default.</p>\n<p>Over time, “frontend engineers should not touch the backend” and “backend engineers should not write pages” became an almost instinctive organizational assumption.</p>\n<p>That assumption is becoming more questionable in the AI era.</p>\n<h2>AI Needs Complete Context</h2>\n<p>AI does not only change a framework detail. It changes a developer’s ability to handle context. The quality of AI-generated code depends heavily on context completeness.</p>\n<p>If a feature’s page, data structure, permission logic, submission flow, error handling, cache behavior, and tests all live in one project, one type system, and related file structures, AI can understand the causal relationships more easily and make more coherent changes.</p>\n<p><img src=\"/assets/posts/2026/frontend-backend-context/03-ai-complete-context.jpg\" alt=\"A complete development context converging into a glowing central model, with pages, data, tests, and rules connected around it\"></p>\n<p>The opposite is also true. If the frontend lives in one repository and the backend in another; if the frontend only sees an API document while the backend cannot see how the page actually uses the data; if deployment is separated; if types have to be synchronized through generated code or documents, then both humans and AI have to reconstruct context from fragments.</p>\n<p>That cost used to be paid by humans. Now it also directly affects the quality of AI output.</p>\n<p>This is why I have started to understand isomorphic full-stack frameworks differently.</p>\n<h2>Isomorphic Frameworks Are Not Only About SSR</h2>\n<p>The value of frameworks such as Next.js is not only that SSR can improve first load experience. It is also not only that API routes and pages can live in the same project. The more important point is that they merge the context of a product feature back together.</p>\n<p>How the page is rendered, how data is loaded, how actions are submitted, where permissions are checked, how types flow, and when caches are invalidated can be handled inside one engineering model. For developers, this reduces cognitive overhead. For AI, it improves context quality.</p>\n<p>I also like the direction represented by Vinext: preserving the full-stack isomorphic development experience associated with Next.js while exploring a more open build and runtime path through ecosystems such as Vite and Cloudflare Workers. Vinext is still experimental, and the framework itself is not the point. The more important trend is that development context is being merged again, and full-stack frameworks will increasingly evolve around AI-friendly engineering models.</p>\n<p>Traditional full-stack frameworks such as Django, Rails, and Laravel are also part of the full-stack path. They have long proved that completing product features in one project is not an outdated idea. The difference is that for modern Web applications with heavier frontend interaction and stronger dependence on component ecosystems, frameworks such as Next.js and Vinext are closer to how current frontend engineering works.</p>\n<h2>Contracts Do Not Disappear</h2>\n<p>One common argument for frontend-backend separation is interface contracts.</p>\n<p>Contracts are still important, but they do not have to take the form of an internal HTTP API pretending to be a public service. External systems, mobile clients, multi-client reuse, and third-party integrations certainly need stable APIs. But data interfaces consumed only by a Web UI do not naturally need to become public contracts.</p>\n<p>After moving to product-layer full stack, contracts do not disappear. They become internalized. They can be TypeScript types, schemas, server actions, component props, database models, unit tests, integration tests, and end-to-end tests. Compared with an API document maintained across separate frontend and backend repositories, these contracts are closer to the code that actually runs and easier for AI and tools to understand together.</p>\n<p>Databases follow a similar logic. In a small project, it is normal for the product layer to read and write the database directly. In a large project, data can be accessed through domain services, storage services, or platform capabilities. But a page-only API should not be created merely to satisfy the doctrine of frontend-backend separation.</p>\n<h2>The Role of BFF Also Changes</h2>\n<p>BFF has a similar issue.</p>\n<p>Many BFF implementations are glue created after frontend-backend separation: forwarding APIs, trimming fields, assembling data, adding some authentication, and converting formats. Their existence proves the opposite of what pure API thinking assumes: a generic API does not always serve page experience well.</p>\n<p>If that is the case, many low-value BFF layers should be absorbed back into product-layer full stack.</p>\n<p>BFF is not always worthless. If it handles complex aggregation, security governance, traffic control, caching, canary rollout, fallback behavior, or shared experience orchestration across multiple products, it can justify being an independent system. But if it exists only so that the frontend never has to touch the backend, it may simply be extra complexity produced by an organizational boundary.</p>\n<h2>Teams Should Be Organized by Business Domain</h2>\n<p>A better organization model is also not to split people by frontend and backend. It is to organize small full-stack teams by business domain.</p>\n<p>Engineers inside one business domain can still have specialties. Some people may be better at interaction, some at data, some at infrastructure. But the team as a whole should own the complete loop from product requirements to production operation.</p>\n<p><img src=\"/assets/posts/2026/frontend-backend-context/04-business-domain-team.jpg\" alt=\"Several business domains arranged around a shared center, each owned by a small full-stack team\"></p>\n<p>When people are split by technical stack for too long, each engineer understands only one segment of the chain. Eventually, nobody fully understands how user needs become working product behavior.</p>\n<p>This does not mean every project should avoid separation. Multi-client reuse of the same API, open platforms, complex core business systems, strong security or compliance requirements, extreme performance constraints, very large-team collaboration, and top-tier consumer applications at major companies can all justify clear frontend-backend boundaries. Multi-client scenarios are especially important: when the same business capability has to serve Web, iOS, Android, Mini Programs, and third-party systems, stable APIs become much more valuable.</p>\n<p>But those should be reasons for separation, not the starting point.</p>\n<h2>A Small Personal Observation</h2>\n<p>My own experience is straightforward. In personal projects, I used to use NestJS plus a Web page. That combination works, but interface definitions, type synchronization, integration work, deployment, file hopping, and context switching keep appearing.</p>\n<p>After moving to full-stack isomorphic solutions such as Vinext, the most obvious change was not writing a few fewer lines of code. It was that one feature could finally be completed in one context. That matters for humans, and it matters even more for AI.</p>\n<p>This is not a rigorous conclusion from a massive production system. It is closer to an architectural preference accumulated while actually writing code. But many architecture decisions eventually return to a plain question: does this structure help people understand the business faster, or does it create more boundaries that have to be synchronized?</p>\n<h2>Default to Full Stack Unless There Is a Reason to Split</h2>\n<p>So my current default judgment is:</p>\n<ul>\n<li>If there is only one primary Web client, prefer product-layer full stack.</li>\n<li>If page data strongly depends on UI shape, prefer product-layer full stack.</li>\n<li>If an API mainly serves its own pages rather than outside consumers, prefer product-layer full stack.</li>\n<li>If the team is not large enough to require strong boundaries, prefer product-layer full stack.</li>\n<li>If the business becomes complex, split by business domain and platform capability before splitting by frontend and backend.</li>\n</ul>\n<p>When discussing architecture, the first question should no longer be:</p>\n<p>Should we separate frontend and backend?</p>\n<p>It should be:</p>\n<p>Does this business have a strong enough reason to split the product context apart?</p>\n<p>If not, frontend-backend separation should no longer be the default. It is not a symbol of advanced engineering. It is a form of complexity that needs to justify itself.</p>\n","date_published":"2026-05-04T00:00:00.000Z","tags":["AI","Full Stack","Frontend-Backend Separation","Architecture"],"language":"en"},{"id":"https://www.lihuanyu.com/en/posts/2025/rethinking-docker-development-linux-redis/","url":"https://www.lihuanyu.com/en/posts/2025/rethinking-docker-development-linux-redis/","title":"Docker Performance: Linux vs Docker Desktop Overhead","summary":"Why Docker feels heavy on macOS and Windows but runs leaner on Linux, including the real costs of filesystems, networking, memory, and container limits.","content_html":"<p>My first impression of Docker came from development machines. On macOS and Windows, starting Docker Desktop meant more memory usage, slower file operations, and the occasional hot-reload problem. It was easy to carry that impression over to Linux servers and assume that Docker itself was heavy.</p>\n<p>That assumption made me hesitate before putting Redis on a small Linux server. The machine had limited CPU and memory, so wasting resources on an extra abstraction layer seemed like a bad trade.</p>\n<p>The mistake was treating Docker Desktop and Docker Engine on Linux as the same runtime. They are not.</p>\n<p><a href=\"/posts/2025/%E9%87%8D%E6%96%B0%E8%AE%A4%E8%AF%86Docker%E7%9A%84%E6%80%A7%E8%83%BD%E5%BC%80%E9%94%80/\">Chinese version of this article</a></p>\n<h2>The short answer</h2>\n<p>Docker adds overhead, but the size and location of that overhead depend on the host and workload.</p>\n<div class=\"table-scroll\"><table>\n<thead>\n<tr>\n<th>Environment</th>\n<th>What runs underneath</th>\n<th>Typical source of overhead</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Docker Engine on Linux</td>\n<td>Containers share the host Linux kernel</td>\n<td>Storage drivers, bridge networking, logging, and resource controls</td>\n</tr>\n<tr>\n<td>Docker Desktop on macOS</td>\n<td>Linux containers run inside a Linux virtual machine</td>\n<td>The VM, file sharing, and host-to-VM filesystem boundaries</td>\n</tr>\n<tr>\n<td>Docker Desktop on Windows</td>\n<td>Linux containers commonly run through WSL2</td>\n<td>The Linux VM boundary and Windows/Linux filesystem access</td>\n</tr>\n<tr>\n<td>A full VM per service</td>\n<td>Each VM includes its own guest kernel</td>\n<td>Guest OS memory, virtual devices, and VM management</td>\n</tr>\n</tbody>\n</table>\n</div><p>For CPU- and memory-bound services on Linux, container performance can be close to running the process directly on the host. Storage-heavy workloads, high packet-rate networking, bind mounts across operating systems, and badly configured logging are more likely to expose measurable costs.</p>\n<p>So the useful question is not simply “Is Docker slow?” It is “Which boundary does this workload cross?”</p>\n<h2>What Docker solves in development</h2>\n<p>In 2017, I used Docker on a small Spring Boot demo. The project needed a specific JDK, Maven, MySQL, initialization data, and a known set of credentials. Moving it to another computer meant repeating setup instructions and hoping nothing had been missed.</p>\n<p>The goal was straightforward: after cloning the project, another developer should be able to start the backend and database with one Compose command.</p>\n<p>That idea still holds up. Docker is particularly useful for moving these dependencies out of a developer’s personal machine:</p>\n<ul>\n<li>Databases such as MySQL, PostgreSQL, and Redis.</li>\n<li>Middleware such as queues, search engines, and object storage emulators.</li>\n<li>Backend services with fixed system dependencies.</li>\n<li>Network relationships between services.</li>\n<li>Initialization scripts and test data.</li>\n</ul>\n<p><img src=\"/assets/legacy/_posts/%E4%BD%BF%E7%94%A8Docker%E8%A7%A3%E5%86%B3%E5%BC%80%E5%8F%91%E7%8E%AF%E5%A2%83%E9%97%AE%E9%A2%98/result1.png\" alt=\"Starting a development environment with Docker Compose\"></p>\n<p>Here Docker solves reproducibility. It turns environment knowledge that used to be passed around verbally into configuration committed to the repository.</p>\n<h2>Why Docker Desktop can feel heavy</h2>\n<p>Linux containers need Linux kernel features. macOS does not provide that kernel, and Windows commonly supplies it through WSL2, so Docker Desktop runs a Linux environment behind the scenes.</p>\n<p>That extra layer is most visible in three places:</p>\n<ul>\n<li>The Linux virtual machine needs its own memory and CPU allocation.</li>\n<li>Bind-mounted source code crosses a host-to-VM filesystem boundary.</li>\n<li>File change notifications may not behave like native Linux events.</li>\n</ul>\n<p>I hit the third problem with a React project on Docker for Windows. The project lived on the Windows filesystem and was mounted into a container. Editing a file changed its contents inside the container, but webpack did not reliably receive the event needed to rebuild the application.</p>\n<p>Polling made the watcher work, but it increased CPU usage. The better default today is to reduce the number of boundaries:</p>\n<ul>\n<li>On Windows, keep active source code inside the WSL2 Linux filesystem.</li>\n<li>On macOS, avoid bind-mounting dependency directories with thousands of small files when a named volume will work.</li>\n<li>Containerize databases and middleware, but keep a pure frontend toolchain on the host when that gives faster feedback.</li>\n<li>Use polling only when native file events cannot cross the boundary reliably.</li>\n</ul>\n<p>Docker Desktop is convenient, but convenience has a resource cost. That cost should not be used as a direct estimate for Docker Engine on a Linux server.</p>\n<h2>Why Docker Engine is lighter on Linux</h2>\n<p>On Linux, containers are processes on the host. They do not boot a separate guest kernel for every service. Isolation mainly comes from kernel features:</p>\n<ul>\n<li>namespaces give a process its own view of processes, networking, mounts, and hostnames;</li>\n<li>cgroups account for and limit CPU, memory, and I/O;</li>\n<li>overlay filesystems combine read-only image layers with a writable container layer.</li>\n</ul>\n<p>This is why CPU and memory performance can be close to native Linux processes. An IBM Research comparison of containers and virtual machines found near-native results for many CPU, memory, and network tests, while also showing that storage and networking choices can matter.</p>\n<p>That benchmark is evidence about particular workloads, not a promise that every container is free. A database writing through overlayfs, a service sending traffic through several network layers, and a CPU-only worker will not have the same profile.</p>\n<h2>Where Linux container overhead appears</h2>\n<p>Docker Engine still has costs. They are usually concrete enough to locate.</p>\n<h3>Storage</h3>\n<p>The writable container layer uses a storage driver such as overlay2. It is convenient for ephemeral files, but a write-heavy database should use a volume or bind mount with a deliberate backup plan. Image layers, stopped containers, and build caches also consume disk space over time.</p>\n<h3>Networking</h3>\n<p>Bridge networking, NAT, and published ports add work to the network path. For most small web services this is not the bottleneck, but latency-sensitive or high-throughput systems should be measured with their real topology.</p>\n<h3>Memory and CPU limits</h3>\n<p>Containers do not receive a useful memory limit automatically. Without one, a process can pressure the whole host just as a native process can. Limits also need headroom: setting them too close to normal usage creates throttling or out-of-memory restarts that look like application instability.</p>\n<h3>Logging</h3>\n<p>The default JSON log driver writes container output to disk. A noisy service without log rotation can fill a small server even when its application data is tiny.</p>\n<h3>Startup and image maintenance</h3>\n<p>Pulling images, unpacking layers, and starting through an entrypoint add work that a long-running service barely notices but a short-lived command may. Images also need security updates and version planning.</p>\n<h2>How to measure Docker overhead on your server</h2>\n<p>A single idle-memory screenshot is not a benchmark. Start with the host and container signals:</p>\n<pre><code class=\"language-bash\">docker stats --no-stream\ndocker system df\nfree -h\ndf -h\n</code></pre>\n<p>Then compare the application under the same workload:</p>\n<ol>\n<li>Use the same binary, dataset, request pattern, and concurrency.</li>\n<li>Warm caches before measuring steady-state latency.</li>\n<li>Separate startup time from long-running throughput.</li>\n<li>Test the actual storage path: writable layer, volume, or bind mount.</li>\n<li>Record CPU, memory, latency percentiles, disk I/O, and errors.</li>\n<li>Repeat the test instead of trusting one run.</li>\n</ol>\n<p>For a container with a memory limit, confirm what Docker applied:</p>\n<pre><code class=\"language-bash\">docker inspect --format '{{.HostConfig.Memory}}' &lt;container-name&gt;\n</code></pre>\n<p>The value is reported in bytes. A result of <code>0</code> means no explicit Docker memory limit is set.</p>\n<p>Measurement also protects against the opposite mistake: blaming Docker for work actually caused by the application, database queries, DNS, or a remote dependency.</p>\n<h2>A Redis container on a small Linux server</h2>\n<p>Redis was the case that changed my own judgment. An idle Redis container on my lightweight server used only a few to a dozen megabytes of memory, with almost no CPU activity. The exact number depends on the Redis version, data, persistence settings, and host, but it was nowhere near the Docker Desktop footprint I had expected.</p>\n<p>The container still needed explicit boundaries: a localhost-only port, persistent data, a Redis memory policy, a container memory limit, log rotation, authentication, and backups.</p>\n<p>The complete configuration is in <a href=\"/en/posts/2026/run-redis-with-docker-compose-on-ubuntu/\">How to Run Redis with Docker Compose on Ubuntu</a>. Keeping that deployment procedure separate makes the performance question easier to answer without burying the operational details.</p>\n<h2>When Docker is a good fit</h2>\n<p>Docker is a good fit when:</p>\n<ul>\n<li>A development environment needs repeatable databases, middleware, or backend dependencies.</li>\n<li>A Linux server needs lightweight services without installing every dependency on the host.</li>\n<li>Runtime versions, networks, volumes, and startup behavior should live in versioned configuration.</li>\n<li>A service should be easy to move to another Linux machine.</li>\n</ul>\n<p>Docker needs more caution when:</p>\n<ul>\n<li>A frontend project on macOS or Windows depends heavily on file watching and small-file I/O.</li>\n<li>A database has heavy writes and needs careful storage and recovery planning.</li>\n<li>A very small server is already close to its memory limit.</li>\n<li>A container is started with no limits, log policy, persistence, or upgrade plan.</li>\n<li>Containers are treated as an impenetrable security boundary from the host.</li>\n</ul>\n<h2>Practical defaults</h2>\n<p>For personal servers and small services, my defaults are:</p>\n<ul>\n<li>Install Docker Engine directly on Linux servers.</li>\n<li>Use Compose instead of keeping long <code>docker run</code> commands in shell history.</li>\n<li>Pin an image major version instead of depending on <code>latest</code> forever.</li>\n<li>Put persistent data in an explicit volume or host directory.</li>\n<li>Add memory limits, restart policies, and log rotation.</li>\n<li>Bind private service ports to <code>127.0.0.1</code> or an internal network.</li>\n<li>Back up important data outside the running container.</li>\n<li>Check release notes and keep a rollback path before upgrades.</li>\n</ul>\n<p>Docker does not remove operational work. It makes more of that work visible as configuration.</p>\n<h2>Conclusion</h2>\n<p>Docker felt heavy to me because I first experienced it through Docker Desktop. On macOS and Windows, the Linux VM and filesystem boundary are real costs. On a Linux server, Docker Engine has a different shape: containers share the host kernel, and CPU or memory overhead is often small.</p>\n<p>The remaining costs are workload-specific. Storage drivers, networking, logging, resource limits, and image maintenance still deserve attention. Measure those boundaries with the real application instead of treating “Docker overhead” as one universal number.</p>\n<h2>Further reading</h2>\n<ul>\n<li><a href=\"/en/posts/2026/run-redis-with-docker-compose-on-ubuntu/\">How to Run Redis with Docker Compose on Ubuntu</a></li>\n<li><a href=\"https://docs.docker.com/engine/containers/run/\">Docker Docs: Running containers</a></li>\n<li><a href=\"https://docs.docker.com/desktop/features/wsl/\">Docker Docs: Docker Desktop WSL 2 backend on Windows</a></li>\n<li><a href=\"https://docs.docker.com/engine/containers/resource_constraints/\">Docker Docs: Resource constraints</a></li>\n<li><a href=\"https://docs.docker.com/engine/network/packet-filtering-firewalls/\">Docker Docs: Packet filtering and firewalls</a></li>\n<li><a href=\"https://research.ibm.com/publications/an-updated-performance-comparison-of-virtual-machines-and-linux-containers\">IBM Research: An Updated Performance Comparison of Virtual Machines and Linux Containers</a></li>\n</ul>\n","date_published":"2025-11-08T00:00:00.000Z","date_modified":"2026-08-01T00:00:00.000Z","tags":["Docker","Linux","Docker Desktop","Containers","Performance","Development Environment"],"language":"en"},{"id":"https://www.lihuanyu.com/posts/2025/%E9%87%8D%E6%96%B0%E8%AE%A4%E8%AF%86Docker%E7%9A%84%E6%80%A7%E8%83%BD%E5%BC%80%E9%94%80/","url":"https://www.lihuanyu.com/posts/2025/%E9%87%8D%E6%96%B0%E8%AE%A4%E8%AF%86Docker%E7%9A%84%E6%80%A7%E8%83%BD%E5%BC%80%E9%94%80/","title":"重新认识 Docker：开发环境、Linux 性能开销与 Redis 实战","summary":"从早期用 Docker 统一开发环境，到后来在 Linux 服务器上部署 Redis，重新梳理 Docker 在开发机和服务器上的真实成本、适用边界和实践细节。","content_html":"<p>我最早接触 Docker，是想解决开发环境不一致的问题。老项目依赖低版本 Node、JDK、Maven、MySQL，换一台机器就可能跑不起来。Docker 的吸引力很直接：把项目依赖的运行环境写进配置文件，让别人拉下代码后用一条命令启动。</p>\n<p>后来在 macOS 和 Windows 上长期使用 Docker Desktop，又形成了另一个印象：Docker 很重。启动后风扇转、内存占用上去、文件监听和热更新偶尔还会出问题。这个印象又让我在低配置 Linux 服务器上不敢轻易使用 Docker。</p>\n<p>English versions: <a href=\"/en/posts/2025/rethinking-docker-development-linux-redis/\">Docker Performance: Linux vs Docker Desktop Overhead</a> and <a href=\"/en/posts/2026/run-redis-with-docker-compose-on-ubuntu/\">How to Run Redis with Docker Compose on Ubuntu</a>.</p>\n<p>直到需要在轻量服务器上部署 Redis 做配置同步，我才重新把这两段经验放在一起看。结论是：Docker 的价值和成本必须区分场景讨论。开发机上的 Docker Desktop、Linux 服务器上的 Docker Engine、用 Compose 编排开发环境、用容器跑 Redis，并不是同一个问题。</p>\n<h2>Docker 最适合解决什么开发环境问题</h2>\n<p>2017 年我用 Docker 改过一个 Spring Boot demo。当时的问题很典型：</p>\n<ul>\n<li>项目需要 JDK 1.8 和 Maven。</li>\n<li>后端依赖 MySQL。</li>\n<li>不同开发者的系统不一样。</li>\n<li>只靠 README 让别人手动装环境，失败概率很高。</li>\n</ul>\n<p>那时最朴素的目标是：别人克隆项目后，不需要在本机安装 MySQL，也不需要追问数据库账号密码和初始化脚本，直接 <code>docker-compose up</code> 就能看到接口返回。</p>\n<p>这个方向今天仍然成立。Docker 很适合把这些东西从开发者电脑上剥离出去：</p>\n<ul>\n<li>数据库，例如 MySQL、PostgreSQL、Redis。</li>\n<li>消息队列、搜索引擎、对象存储模拟器等中间件。</li>\n<li>需要固定系统依赖的后端服务。</li>\n<li>多个服务之间的网络关系。</li>\n<li>初始化脚本、测试数据和本地端口映射。</li>\n</ul>\n<p>当时的 demo 用一个 web 容器跑 Spring Boot，用一个 MySQL 容器提供数据库，再由 Compose 统一启动。浏览器打开 <code>localhost:8080</code> 就能看到接口结果。</p>\n<p><img src=\"/assets/legacy/_posts/%E4%BD%BF%E7%94%A8Docker%E8%A7%A3%E5%86%B3%E5%BC%80%E5%8F%91%E7%8E%AF%E5%A2%83%E9%97%AE%E9%A2%98/result1.png\" alt=\"Docker Compose 启动开发环境\"></p>\n<p>这类场景里，Docker 解决的是“环境可复制”。它不只是省掉安装步骤，更重要的是把口口相传的环境知识变成仓库里的配置。</p>\n<h2>开发机上的坑：文件系统和热更新</h2>\n<p>开发环境并不是只要容器能启动就结束了。前端项目还有热更新、文件监听、依赖安装和大量小文件读写。</p>\n<p>我后来在 Docker for Windows 上遇到过一个问题：React 项目放在 Windows 文件系统里，通过 volume 挂到容器内，页面能启动，但编辑文件后容器里的 webpack 不会触发重新编译。文件内容已经同步到容器里，问题出在文件变更通知没有可靠传递。</p>\n<p>那个年代的解决方案很偏 workaround：额外跑一个 watcher，把 Windows 里的文件变动转成容器能感知的变动。它解决了当时的问题，但不是一个今天还值得推荐的默认方案。</p>\n<p>今天更稳妥的判断是：</p>\n<ul>\n<li>Windows 开发尽量使用 WSL2，并把项目放在 Linux 发行版的文件系统里，而不是放在 Windows 盘再挂载进去。</li>\n<li>macOS 上如果遇到大量小文件 I/O 或热更新变慢，要减少 bind mount 的范围，依赖目录尽量用 named volume。</li>\n<li>前端热更新如果必须跨宿主机和容器边界，必要时启用工具自身的 polling 模式，但它会增加 CPU 开销。</li>\n<li>对纯前端项目，不必为了“统一环境”强行把所有开发流程都塞进容器。很多时候本机 Node + 容器中间件更舒服。</li>\n</ul>\n<p>Docker Desktop 的文档也把 Windows 上的 WSL2 工作流作为重要路径，并建议代码放在 Linux 发行版内获得更好的开发体验。这个建议和早期踩坑的方向是一致的。</p>\n<p>所以，开发机上的 Docker 是一把工具，不是宗教。它适合统一数据库、中间件和后端依赖；对高频热更新的前端开发，要根据文件系统表现做取舍。</p>\n<h2>为什么 Linux 服务器上的 Docker 轻很多</h2>\n<p>我以前觉得 Docker 重，主要来自 macOS 和 Windows 上的体验。但这两个系统不能直接运行 Linux 容器，需要 Docker Desktop 在背后准备 Linux 环境。</p>\n<p>在 Windows 上，Docker Desktop 通常通过 WSL2 后端运行；在 macOS 上，也需要一个 Linux 虚拟化环境来承载容器。资源占用和文件系统映射开销，很大一部分来自这层虚拟化和宿主机/虚拟机之间的边界。</p>\n<p>Linux 服务器上的 Docker Engine 则不同。Docker 官方文档对容器的描述很直接：容器是运行在宿主机上的进程，只是拥有自己的文件系统、网络和进程树隔离。实现隔离主要依赖 Linux 内核能力：</p>\n<ul>\n<li>namespace：隔离进程、网络、挂载点、主机名等视图。</li>\n<li>cgroups：限制和统计 CPU、内存、I/O 等资源。</li>\n<li>union filesystem/overlayfs：让镜像层和容器可写层组合起来。</li>\n</ul>\n<p>这意味着在 Linux 上，容器不是一台完整虚拟机。它仍然有开销，但开销通常远小于“每个服务一台虚拟机”的模型。</p>\n<p>IBM Research 的容器性能研究也给过类似结论：在很多 CPU、内存和网络基准测试里，Linux 容器接近裸机表现；明显差异更多出现在特定 I/O、网络路径或存储驱动场景。这个结论不能简单翻译成“Docker 永远无损耗”，但足以说明：把 macOS/Windows 上 Docker Desktop 的体感，直接套到 Linux 服务器上是不准确的。</p>\n<p>更准确的说法是：</p>\n<ul>\n<li>Docker Desktop：开发体验工具，便利性强，但包含虚拟化层和文件系统映射成本。</li>\n<li>Docker Engine on Linux：服务器运行时，直接使用 Linux 内核能力，适合部署轻量服务。</li>\n<li>Docker Desktop for Linux 也会运行 VM，它和服务器上直接安装 Docker Engine 不是一回事。</li>\n</ul>\n<p>这也是我后来敢在轻量服务器上用 Docker 跑 Redis 的原因。</p>\n<h2>Linux 上也不是完全没有成本</h2>\n<p>把 Docker 放到 Linux 服务器上，并不代表可以完全不管资源。</p>\n<p>几个成本仍然存在：</p>\n<ul>\n<li>镜像和容器层会占用磁盘，需要定期清理不用的镜像。</li>\n<li>日志默认可能写到 Docker 管理目录，长时间运行要配置日志轮转。</li>\n<li>bridge 网络和端口映射有一点网络开销。</li>\n<li>overlayfs 对某些写密集型场景不一定是最佳选择。</li>\n<li>bind mount、volume、权限、UID/GID 需要认真处理。</li>\n<li>容器默认不会自动限制内存，服务失控时仍可能拖垮宿主机。</li>\n</ul>\n<p>所以合理做法不是“因为 Docker 很轻就随便跑”，而是给服务加上边界：限制内存、限制日志、持久化数据、明确端口暴露范围。</p>\n<h2>在 Ubuntu 上安装 Docker Engine</h2>\n<p>服务器上建议安装 Docker Engine，而不是 Docker Desktop。Docker 官方文档提供了 Ubuntu 的 apt 仓库安装方式，命令会随版本演进，长期以官方页面为准。</p>\n<p>一组常见步骤如下：</p>\n<pre><code class=\"language-bash\">sudo apt-get update\nsudo apt-get install -y ca-certificates curl\n\nsudo install -m 0755 -d /etc/apt/keyrings\nsudo curl -fsSL https://download.docker.com/linux/ubuntu/gpg -o /etc/apt/keyrings/docker.asc\nsudo chmod a+r /etc/apt/keyrings/docker.asc\n\necho \\\n  &quot;deb [arch=$(dpkg --print-architecture) signed-by=/etc/apt/keyrings/docker.asc] https://download.docker.com/linux/ubuntu \\\n  $(. /etc/os-release &amp;&amp; echo &quot;${UBUNTU_CODENAME:-$VERSION_CODENAME}&quot;) stable&quot; | \\\n  sudo tee /etc/apt/sources.list.d/docker.list &gt; /dev/null\n\nsudo apt-get update\nsudo apt-get install -y docker-ce docker-ce-cli containerd.io docker-buildx-plugin docker-compose-plugin\n</code></pre>\n<p>安装后启动并设置开机自启：</p>\n<pre><code class=\"language-bash\">sudo systemctl enable --now docker\ndocker --version\ndocker compose version\n</code></pre>\n<p>如果不想每次都写 <code>sudo</code>，可以把当前用户加入 <code>docker</code> 组：</p>\n<pre><code class=\"language-bash\">sudo usermod -aG docker &quot;$USER&quot;\n</code></pre>\n<p>这个操作需要重新登录才生效。也要注意，能访问 Docker socket 的用户基本等同于能获得宿主机 root 权限，不应该随便给普通账号开放。</p>\n<h2>用 Compose 跑一个受限制的 Redis</h2>\n<p>我当时的目标是部署一个轻量 Redis，用来做配置同步。Redis 很适合作为 Docker 实战样本：镜像成熟、启动快、资源占用低，同时又涉及端口、持久化、内存限制和安全配置。</p>\n<p>先创建目录：</p>\n<pre><code class=\"language-bash\">mkdir -p ~/services/redis-config/data\ncd ~/services/redis-config\n</code></pre>\n<p>准备 <code>.env</code>：</p>\n<pre><code class=\"language-bash\">REDIS_PASSWORD=change-this-password\n</code></pre>\n<p>准备 <code>compose.yaml</code>：</p>\n<pre><code class=\"language-yaml\">services:\n  redis:\n    image: redis:8-alpine\n    container_name: redis-config\n    restart: unless-stopped\n    ports:\n      - &quot;127.0.0.1:6379:6379&quot;\n    command:\n      - redis-server\n      - --appendonly\n      - &quot;yes&quot;\n      - --maxmemory\n      - &quot;64mb&quot;\n      - --maxmemory-policy\n      - allkeys-lru\n      - --requirepass\n      - &quot;${REDIS_PASSWORD:?set REDIS_PASSWORD}&quot;\n    volumes:\n      - ./data:/data\n    mem_limit: 128m\n</code></pre>\n<p>这里有几个关键选择：</p>\n<ul>\n<li>使用 <code>redis:8-alpine</code>，固定主版本，避免 <code>latest</code> 随时间漂移。</li>\n<li>端口绑定到 <code>127.0.0.1</code>，默认只允许本机访问。</li>\n<li>开启 AOF，把数据写到挂载目录。</li>\n<li>设置 Redis 自身的 <code>maxmemory</code> 和淘汰策略。</li>\n<li>设置容器内存上限，避免 Redis 或异常情况吃掉整台机器。</li>\n<li>使用 <code>restart: unless-stopped</code>，服务器重启后自动恢复。</li>\n</ul>\n<p>如果 Redis 只是同机应用使用，绑定 <code>127.0.0.1</code> 是更稳妥的默认值。如果需要跨服务器访问或做复制，应该绑定内网 IP，并用云安全组只放行对端内网地址。不要把 Redis 直接暴露到公网。</p>\n<p>Docker 官方文档还提醒过一个容易忽略的点：发布容器端口可能绕过宿主机上 <code>ufw</code> 或 <code>firewalld</code> 的部分规则。云服务器上更应该同时依赖安全组、内网地址绑定和服务自身认证，而不是只相信本机防火墙。</p>\n<p>启动：</p>\n<pre><code class=\"language-bash\">docker compose up -d\ndocker compose ps\n</code></pre>\n<p>验证：</p>\n<pre><code class=\"language-bash\">docker compose exec redis redis-cli\n</code></pre>\n<p>进入后执行：</p>\n<pre><code class=\"language-text\">AUTH change-this-password\nPING\n</code></pre>\n<p>返回 <code>PONG</code> 就说明 Redis 正常工作。</p>\n<p>观察资源：</p>\n<pre><code class=\"language-bash\">docker stats redis-config --no-stream\nfree -h\n</code></pre>\n<p>如果要看 Redis 自身的内存统计，可以进入 <code>redis-cli</code> 后执行：</p>\n<pre><code class=\"language-text\">AUTH change-this-password\nINFO memory\n</code></pre>\n<p>在我的轻量服务器上，一个空载 Redis 容器的内存占用只有几 MB 到十几 MB 级别，CPU 基本可以忽略。实际数据会随 Redis 版本、数据量、配置和宿主机环境变化，但这个量级足以说明：低配置 Linux 服务器跑一个轻量 Redis 容器并不夸张。</p>\n<h2>什么时候适合用 Docker</h2>\n<p>这些实践放在一起后，我对 Docker 的判断更清晰了。</p>\n<p>适合用 Docker 的场景：</p>\n<ul>\n<li>需要统一数据库、中间件、后端依赖的开发环境。</li>\n<li>服务器上部署轻量服务，希望减少手工安装和环境污染。</li>\n<li>多个服务需要明确网络关系、启动顺序和环境变量。</li>\n<li>希望通过镜像版本固定运行时。</li>\n<li>希望服务可以快速迁移到另一台 Linux 机器。</li>\n</ul>\n<p>需要谨慎的场景：</p>\n<ul>\n<li>前端项目在 macOS/Windows 上强依赖大量文件监听和热更新。</li>\n<li>数据库写入很重，需要仔细评估磁盘、volume、备份和恢复。</li>\n<li>服务器内存极低，例如 512MB，还要跑多个服务。</li>\n<li>只会 <code>docker run</code>，但没有规划日志、持久化、安全和升级。</li>\n<li>把 Docker 当成安全边界，以为容器里出问题不会影响宿主机。</li>\n</ul>\n<p>Docker 的最佳位置，是把运行环境变成代码，同时给服务加上清晰边界。它不是为了替代所有本机开发工具，也不是为了掩盖运维设计。</p>\n<h2>一套更稳妥的默认实践</h2>\n<p>如果是个人服务器或小项目，我会按下面的方式使用 Docker：</p>\n<ul>\n<li>Linux 服务器安装 Docker Engine，不安装 Docker Desktop。</li>\n<li>使用 <code>docker compose</code> 管理服务，而不是把超长 <code>docker run</code> 命令散落在笔记里。</li>\n<li>镜像固定主版本，例如 <code>redis:8-alpine</code>，不要长期依赖 <code>latest</code>。</li>\n<li>数据写到明确的 volume 或宿主机目录。</li>\n<li>容器设置重启策略和内存上限。</li>\n<li>服务端口默认绑定 <code>127.0.0.1</code> 或内网 IP。</li>\n<li>需要公网访问时，前面放 Nginx/Caddy/网关，不让数据库类服务裸露。</li>\n<li>定期查看 <code>docker ps</code>、<code>docker stats</code>、<code>docker logs</code> 和磁盘占用。</li>\n<li>升级镜像前看 release notes，升级后保留回滚路径。</li>\n<li>对重要数据做宿主机级备份，而不是以为容器还在数据就安全。</li>\n</ul>\n<p>这套做法不复杂，但能避免很多“容器跑起来了，后来不好维护”的问题。</p>\n<h2>总结</h2>\n<p>我对 Docker 的认知变化，基本经历了三个阶段。</p>\n<p>最开始，它是统一开发环境的工具：把 JDK、MySQL、后端服务和网络关系用 Compose 固化下来，减少项目启动成本。</p>\n<p>后来，Docker Desktop 在 macOS/Windows 上的体感让我觉得它很重，尤其是文件系统、热更新和资源占用。</p>\n<p>再后来，在 Linux 服务器上实际跑 Redis，才发现 Docker Engine 的运行成本和 Docker Desktop 的开发机体验不能混为一谈。对轻量服务来说，Linux 上的 Docker 很实用，关键是配好持久化、资源限制和安全边界。</p>\n<p>Docker 不是性能负担的代名词，也不是万能部署答案。它更像一层可复制的运行环境描述。用得克制、边界清楚，就很适合个人项目和小型服务。</p>\n<h2>扩展阅读</h2>\n<ul>\n<li><a href=\"https://docs.docker.com/engine/containers/run/\">Docker Docs: Running containers</a></li>\n<li><a href=\"https://docs.docker.com/engine/install/ubuntu/\">Docker Docs: Install Docker Engine on Ubuntu</a></li>\n<li><a href=\"https://docs.docker.com/desktop/features/wsl/\">Docker Docs: Docker Desktop WSL 2 backend on Windows</a></li>\n<li><a href=\"https://docs.docker.com/engine/network/packet-filtering-firewalls/\">Docker Docs: Packet filtering and firewalls</a></li>\n<li><a href=\"https://hub.docker.com/_/redis/\">Docker Hub: Redis Official Image</a></li>\n<li><a href=\"https://research.ibm.com/publications/an-updated-performance-comparison-of-virtual-machines-and-linux-containers\">IBM Research: An Updated Performance Comparison of Virtual Machines and Linux Containers</a></li>\n</ul>\n","date_published":"2025-11-08T00:00:00.000Z","date_modified":"2026-05-04T00:00:00.000Z","tags":["Docker","Redis","Linux","容器化","运维","开发环境"],"language":"zh"},{"id":"https://www.lihuanyu.com/posts/2025/NestJS-%E6%A1%86%E6%9E%B6%E4%B8%8B%E7%9A%84%E4%B8%89%E7%A7%8D%E6%B5%8B%E8%AF%95%E7%B1%BB%E5%9E%8B%E5%AF%B9%E6%AF%94%E5%88%86%E6%9E%90/","url":"https://www.lihuanyu.com/posts/2025/NestJS-%E6%A1%86%E6%9E%B6%E4%B8%8B%E7%9A%84%E4%B8%89%E7%A7%8D%E6%B5%8B%E8%AF%95%E7%B1%BB%E5%9E%8B%E5%AF%B9%E6%AF%94%E5%88%86%E6%9E%90/","title":"NestJS 框架下的三种测试类型对比分析","summary":"以 NestJS 为例，对比单元测试、集成测试和端到端测试的测试范围、执行特性、实现方式和适用场景。","content_html":"<h2>概述</h2>\n<p>本文以NestJS框架为例，深入对比分析单元测试、集成测试和端到端(E2E)测试的核心区别，帮助开发者在实际项目中选择合适的测试策略。</p>\n<h2>1. 三种测试类型的核心区别</h2>\n<h3>1.1 定义与测试范围对比</h3>\n<div class=\"table-scroll\"><table>\n<thead>\n<tr>\n<th>测试类型</th>\n<th>定义</th>\n<th>测试范围</th>\n<th>NestJS中的体现</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>单元测试</strong></td>\n<td>测试单个函数、方法或类的逻辑</td>\n<td>最小可测试单元</td>\n<td>Service方法、Controller方法、Pipe、Guard等</td>\n</tr>\n<tr>\n<td><strong>集成测试</strong></td>\n<td>测试多个模块间的交互</td>\n<td>模块间接口和数据流</td>\n<td>Service与Repository交互、Module间通信</td>\n</tr>\n<tr>\n<td><strong>E2E测试</strong></td>\n<td>测试完整的用户场景</td>\n<td>整个应用流程</td>\n<td>HTTP请求到响应的完整链路</td>\n</tr>\n</tbody>\n</table>\n</div><h3>1.2 执行特性对比</h3>\n<div class=\"table-scroll\"><table>\n<thead>\n<tr>\n<th>特性</th>\n<th>单元测试</th>\n<th>集成测试</th>\n<th>E2E测试</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>执行速度</strong></td>\n<td>极快(毫秒级)</td>\n<td>中等(秒级)</td>\n<td>较慢(分钟级)</td>\n</tr>\n<tr>\n<td><strong>隔离性</strong></td>\n<td>完全隔离</td>\n<td>部分隔离</td>\n<td>无隔离</td>\n</tr>\n<tr>\n<td><strong>依赖处理</strong></td>\n<td>Mock所有依赖</td>\n<td>Mock外部依赖</td>\n<td>使用真实依赖</td>\n</tr>\n<tr>\n<td><strong>环境要求</strong></td>\n<td>无需外部环境</td>\n<td>需要部分真实环境</td>\n<td>需要完整环境</td>\n</tr>\n</tbody>\n</table>\n</div><h2>2. NestJS框架下的具体实现对比</h2>\n<h3>2.1 单元测试示例</h3>\n<p><strong>测试目标：UserService中的创建用户方法</strong></p>\n<pre><code class=\"language-typescript\">// user.service.ts\n@Injectable()\nexport class UserService {\n  constructor(private userRepository: UserRepository) {}\n\n  async createUser(userData: CreateUserDto): Promise {\n    const existingUser = await this.userRepository.findByEmail(userData.email);\n    if (existingUser) {\n      throw new ConflictException('User already exists');\n    }\n    return this.userRepository.create(userData);\n  }\n}\n</code></pre>\n<p><strong>单元测试实现：</strong></p>\n<pre><code class=\"language-typescript\">// user.service.spec.ts\ndescribe('UserService', () =&gt; {\n  let service: UserService;\n  let mockRepository: jest.Mocked;\n\n  beforeEach(async () =&gt; {\n    const mockRepo = {\n      findByEmail: jest.fn(),\n      create: jest.fn(),\n    };\n\n    const module = await Test.createTestingModule({\n      providers: [\n        UserService,\n        { provide: UserRepository, useValue: mockRepo },\n      ],\n    }).compile();\n\n    service = module.get(UserService);\n    mockRepository = module.get(UserRepository);\n  });\n\n  it('should create user when email not exists', async () =&gt; {\n    // Arrange\n    const userData = { email: 'test@example.com', name: 'Test User' };\n    mockRepository.findByEmail.mockResolvedValue(null);\n    mockRepository.create.mockResolvedValue({ id: 1, ...userData });\n\n    // Act\n    const result = await service.createUser(userData);\n\n    // Assert\n    expect(mockRepository.findByEmail).toHaveBeenCalledWith(userData.email);\n    expect(mockRepository.create).toHaveBeenCalledWith(userData);\n    expect(result).toEqual({ id: 1, ...userData });\n  });\n});\n</code></pre>\n<p><strong>特点分析：</strong></p>\n<ul>\n<li>\n<p>✅ <strong>完全隔离</strong>：Mock了UserRepository依赖</p>\n</li>\n<li>\n<p>✅ <strong>快速执行</strong>：无需数据库连接</p>\n</li>\n<li>\n<p>✅ <strong>精确验证</strong>：只测试业务逻辑</p>\n</li>\n<li>\n<p>❌ <strong>无法发现</strong>：Repository接口变更问题</p>\n</li>\n</ul>\n<h3>2.2 集成测试示例</h3>\n<p><strong>测试目标：UserService与真实数据库的交互</strong></p>\n<pre><code class=\"language-typescript\">// user.integration.spec.ts\ndescribe('UserService Integration', () =&gt; {\n  let app: INestApplication;\n  let service: UserService;\n  let repository: Repository;\n\n  beforeAll(async () =&gt; {\n    const module = await Test.createTestingModule({\n      imports: [\n        TypeOrmModule.forRoot({\n          type: 'sqlite',\n          database: ':memory:',\n          entities: [User],\n          synchronize: true,\n        }),\n        TypeOrmModule.forFeature([User]),\n      ],\n      providers: [UserService, UserRepository],\n    }).compile();\n\n    app = module.createNestApplication();\n    await app.init();\n    \n    service = module.get(UserService);\n    repository = module.get&gt;(getRepositoryToken(User));\n  });\n\n  beforeEach(async () =&gt; {\n    await repository.clear(); // 清理测试数据\n  });\n\n  it('should create user and save to database', async () =&gt; {\n    // Arrange\n    const userData = { email: 'test@example.com', name: 'Test User' };\n\n    // Act\n    const createdUser = await service.createUser(userData);\n\n    // Assert\n    expect(createdUser.id).toBeDefined();\n    expect(createdUser.email).toBe(userData.email);\n    \n    // 验证数据确实保存到数据库\n    const savedUser = await repository.findOne({ where: { email: userData.email } });\n    expect(savedUser).toBeTruthy();\n  });\n\n  it('should throw conflict when user already exists', async () =&gt; {\n    // Arrange\n    const userData = { email: 'test@example.com', name: 'Test User' };\n    await repository.save(userData); // 预先创建用户\n\n    // Act &amp; Assert\n    await expect(service.createUser(userData)).rejects.toThrow(ConflictException);\n  });\n});\n</code></pre>\n<p><strong>特点分析：</strong></p>\n<ul>\n<li>\n<p>✅ <strong>真实交互</strong>：使用真实数据库操作</p>\n</li>\n<li>\n<p>✅ <strong>接口验证</strong>：能发现Service与Repository接口问题</p>\n</li>\n<li>\n<p>✅ <strong>数据验证</strong>：确认数据正确保存</p>\n</li>\n<li>\n<p>❌ <strong>执行较慢</strong>：需要数据库操作</p>\n</li>\n<li>\n<p>❌ <strong>环境依赖</strong>：需要配置测试数据库</p>\n</li>\n</ul>\n<h3>2.3 E2E测试示例</h3>\n<p><strong>测试目标：完整的用户注册API流程</strong></p>\n<pre><code class=\"language-typescript\">// user.e2e-spec.ts\ndescribe('User E2E', () =&gt; {\n  let app: INestApplication;\n  let httpServer: any;\n\n  beforeAll(async () =&gt; {\n    const module = await Test.createTestingModule({\n      imports: [AppModule], // 导入完整应用模块\n    }).compile();\n\n    app = module.createNestApplication();\n    app.useGlobalPipes(new ValidationPipe()); // 应用全局管道\n    await app.init();\n    \n    httpServer = app.getHttpServer();\n  });\n\n  beforeEach(async () =&gt; {\n    // 清理测试数据\n    const userRepository = app.get&gt;(getRepositoryToken(User));\n    await userRepository.clear();\n  });\n\n  it('/users (POST) - should create user successfully', async () =&gt; {\n    // Arrange\n    const userData = {\n      email: 'test@example.com',\n      name: 'Test User',\n      password: 'password123'\n    };\n\n    // Act\n    const response = await request(httpServer)\n      .post('/users')\n      .send(userData)\n      .expect(201);\n\n    // Assert\n    expect(response.body).toMatchObject({\n      id: expect.any(Number),\n      email: userData.email,\n      name: userData.name,\n    });\n    expect(response.body.password).toBeUndefined(); // 密码不应返回\n  });\n\n  it('/users (POST) - should return 409 when user exists', async () =&gt; {\n    // Arrange\n    const userData = {\n      email: 'test@example.com',\n      name: 'Test User',\n      password: 'password123'\n    };\n\n    // 先创建用户\n    await request(httpServer)\n      .post('/users')\n      .send(userData)\n      .expect(201);\n\n    // Act &amp; Assert\n    await request(httpServer)\n      .post('/users')\n      .send(userData)\n      .expect(409);\n  });\n\n  it('/users (POST) - should validate input data', async () =&gt; {\n    // Act &amp; Assert\n    await request(httpServer)\n      .post('/users')\n      .send({\n        email: 'invalid-email', // 无效邮箱\n        name: '', // 空名称\n      })\n      .expect(400);\n  });\n});\n</code></pre>\n<p><strong>特点分析：</strong></p>\n<ul>\n<li>\n<p>✅ <strong>完整流程</strong>：测试HTTP请求到数据库的完整链路</p>\n</li>\n<li>\n<p>✅ <strong>真实场景</strong>：模拟用户实际操作</p>\n</li>\n<li>\n<p>✅ <strong>全面验证</strong>：包含验证、异常处理、响应格式等</p>\n</li>\n<li>\n<p>❌ <strong>执行最慢</strong>：启动完整应用</p>\n</li>\n<li>\n<p>❌ <strong>维护成本高</strong>：接口变更需要同步更新</p>\n</li>\n</ul>\n<h2>3. 在NestJS项目中的选择策略</h2>\n<h3>3.1 测试金字塔在NestJS中的应用</h3>\n<pre><code class=\"language-plaintext\">        E2E Tests (10%)\n      ┌─────────────────┐\n      │   核心业务流程   │\n      └─────────────────┘\n    \n    Integration Tests (20%)\n   ┌───────────────────────┐\n   │  Service ↔ Repository │\n   │  Module间交互          │\n   └───────────────────────┘\n\nUnit Tests (70%)\n┌─────────────────────────────┐\n│ Service方法、Controller方法  │\n│ Pipe、Guard、Interceptor    │\n│ 工具函数、业务逻辑           │\n└─────────────────────────────┘\n</code></pre>\n<h3>3.2 具体应用建议</h3>\n<p><strong>单元测试重点关注：</strong></p>\n<ul>\n<li>\n<p>Service中的业务逻辑方法</p>\n</li>\n<li>\n<p>Controller中的参数处理和响应格式化</p>\n</li>\n<li>\n<p>自定义Pipe的数据转换逻辑</p>\n</li>\n<li>\n<p>Guard的权限验证逻辑</p>\n</li>\n<li>\n<p>工具函数和算法</p>\n</li>\n</ul>\n<p><strong>集成测试重点关注：</strong></p>\n<ul>\n<li>\n<p>Service与Repository的数据操作</p>\n</li>\n<li>\n<p>Module间的依赖注入</p>\n</li>\n<li>\n<p>第三方服务的集成（如Redis、消息队列）</p>\n</li>\n<li>\n<p>数据库事务处理</p>\n</li>\n</ul>\n<p><strong>E2E测试重点关注：</strong></p>\n<ul>\n<li>\n<p>用户注册/登录流程</p>\n</li>\n<li>\n<p>核心业务操作流程</p>\n</li>\n<li>\n<p>权限控制的完整验证</p>\n</li>\n<li>\n<p>错误处理的用户体验</p>\n</li>\n</ul>\n<h2>4. 实际项目中的测试配置</h2>\n<h3>4.1 package.json测试脚本</h3>\n<pre><code class=\"language-json\">{\n  &quot;scripts&quot;: {\n    &quot;test&quot;: &quot;jest&quot;,\n    &quot;test:watch&quot;: &quot;jest --watch&quot;,\n    &quot;test:cov&quot;: &quot;jest --coverage&quot;,\n    &quot;test:integration&quot;: &quot;jest --config ./test/jest-integration.json&quot;,\n    &quot;test:e2e&quot;: &quot;jest --config ./test/jest-e2e.json&quot;\n  }\n}\n</code></pre>\n<h3>4.2 Jest配置文件</h3>\n<p><strong>单元测试配置 (jest.config.js):</strong></p>\n<pre><code class=\"language-javascript\">module.exports = {\n  moduleFileExtensions: ['js', 'json', 'ts'],\n  rootDir: 'src',\n  testRegex: '.*\\\\.spec\\\\.ts$', // 只匹配 .spec.ts 文件\n  transform: { '^.+\\\\.(t|j)s$': 'ts-jest' },\n  collectCoverageFrom: ['**/*.(t|j)s'],\n  coverageDirectory: '../coverage',\n  testEnvironment: 'node',\n  testPathIgnorePatterns: ['.*\\\\.integration\\\\.spec\\\\.ts$'], // 排除集成测试\n};\n</code></pre>\n<p><strong>集成测试配置 (test/jest-integration.json):</strong></p>\n<pre><code class=\"language-json\">{\n  &quot;moduleFileExtensions&quot;: [&quot;js&quot;, &quot;json&quot;, &quot;ts&quot;],\n  &quot;rootDir&quot;: &quot;../src&quot;,\n  &quot;testEnvironment&quot;: &quot;node&quot;,\n  &quot;testRegex&quot;: &quot;.*\\\\.integration\\\\.spec\\\\.ts$&quot;,\n  &quot;transform&quot;: { &quot;^.+\\\\.(t|j)s$&quot;: &quot;ts-jest&quot; },\n  &quot;setupFilesAfterEnv&quot;: [&quot;/../test/integration-setup.ts&quot;]\n}\n</code></pre>\n<p><strong>E2E测试配置 (test/jest-e2e.json):</strong></p>\n<pre><code class=\"language-json\">{\n  &quot;moduleFileExtensions&quot;: [&quot;js&quot;, &quot;json&quot;, &quot;ts&quot;],\n  &quot;rootDir&quot;: &quot;.&quot;,\n  &quot;testEnvironment&quot;: &quot;node&quot;,\n  &quot;testRegex&quot;: &quot;.e2e-spec.ts$&quot;,\n  &quot;transform&quot;: { &quot;^.+\\\\.(t|j)s$&quot;: &quot;ts-jest&quot; }\n}\n</code></pre>\n<p><strong>集成测试环境设置 (test/integration-setup.ts):</strong></p>\n<pre><code class=\"language-typescript\">import { Test } from '@nestjs/testing';\nimport { TypeOrmModule } from '@nestjs/typeorm';\n\n// 全局集成测试配置\nbeforeAll(async () =&gt; {\n  // 设置测试数据库连接等\n});\n\nafterAll(async () =&gt; {\n  // 清理资源\n});\n</code></pre>\n<h2>5. 总结</h2>\n<p>在NestJS框架下，三种测试类型各有其适用场景：</p>\n<p><strong>选择单元测试当：</strong></p>\n<ul>\n<li>\n<p>验证复杂业务逻辑</p>\n</li>\n<li>\n<p>需要快速反馈</p>\n</li>\n<li>\n<p>测试覆盖率要求高</p>\n</li>\n</ul>\n<p><strong>选择集成测试当：</strong></p>\n<ul>\n<li>\n<p>验证数据库操作</p>\n</li>\n<li>\n<p>测试模块间交互</p>\n</li>\n<li>\n<p>确保接口契约正确</p>\n</li>\n</ul>\n<p><strong>选择E2E测试当：</strong></p>\n<ul>\n<li>\n<p>验证关键业务流程</p>\n</li>\n<li>\n<p>确保用户体验</p>\n</li>\n<li>\n<p>发布前的最终验证</p>\n</li>\n</ul>\n<p>合理的测试策略应该是70%单元测试 + 20%集成测试 + 10%E2E测试，这样既能保证代码质量，又能控制测试维护成本。</p>\n<h2>6. 扩展：其他测试类型的补充说明</h2>\n<h3>6.1 测试分类的两个维度</h3>\n<p>虽然本文重点讨论单元测试、集成测试和E2E测试，但在实际项目中还存在其他测试类型。理解测试的分类维度很重要：</p>\n<p><strong>按执行方式分类：</strong></p>\n<ul>\n<li>\n<p><strong>自动化测试</strong>：通过代码自动执行（本文重点）</p>\n</li>\n<li>\n<p><strong>工具驱动测试</strong>：使用专门工具执行</p>\n</li>\n<li>\n<p><strong>手工测试</strong>：需要人工操作</p>\n</li>\n</ul>\n<p><strong>按测试目标分类：</strong></p>\n<ul>\n<li>\n<p><strong>功能性测试</strong>：验证功能是否正确实现</p>\n</li>\n<li>\n<p><strong>非功能性测试</strong>：验证性能、安全、可用性等</p>\n</li>\n</ul>\n<h3>6.2 代码驱动的自动化测试 vs 其他测试类型</h3>\n<h4>本文讨论的三种测试（代码驱动）</h4>\n<pre><code class=\"language-typescript\">// 完全通过代码自动执行\ndescribe('UserService', () =&gt; {\n  it('should create user when email not exists', async () =&gt; {\n    // 自动化的断言检查\n    expect(result.email).toBe('test@example.com');\n    expect(mockRepository.create).toHaveBeenCalledWith(userData);\n  });\n});\n</code></pre>\n<p><strong>特点：</strong></p>\n<ul>\n<li>\n<p>✅ 完全自动化执行</p>\n</li>\n<li>\n<p>✅ 可集成到CI/CD流程</p>\n</li>\n<li>\n<p>✅ 开发过程中持续运行</p>\n</li>\n<li>\n<p>✅ 快速反馈和问题定位</p>\n</li>\n</ul>\n<h4>其他测试类型（工具/人工驱动）</h4>\n<p><strong>API测试（工具驱动）：</strong></p>\n<pre><code class=\"language-bash\"># 使用Postman/Newman\nnewman run api-tests.postman_collection.json\n\n# 使用专门的API测试工具\ncurl -X POST http://localhost:3000/users \\\n  -H &quot;Content-Type: application/json&quot; \\\n  -d '{&quot;email&quot;:&quot;test@example.com&quot;,&quot;name&quot;:&quot;Test User&quot;}'\n</code></pre>\n<p><strong>性能测试（工具驱动）：</strong></p>\n<pre><code class=\"language-bash\"># 使用Artillery进行负载测试\nartillery run load-test.yml\n\n# 使用JMeter\njmeter -n -t performance-test.jmx\n</code></pre>\n<p><strong>安全测试（工具驱动）：</strong></p>\n<pre><code class=\"language-bash\"># 依赖漏洞扫描\nnpm audit\nsnyk test\n\n# 代码安全扫描\neslint --ext .ts src/ --config .eslintrc-security.js\n</code></pre>\n<h3>6.3 完整的测试策略配置</h3>\n<h4>package.json中的完整测试脚本</h4>\n<pre><code class=\"language-json\">{\n  &quot;scripts&quot;: {\n    // 代码驱动的自动化测试（本文重点）\n    &quot;test&quot;: &quot;jest&quot;,\n    &quot;test:unit&quot;: &quot;jest --config ./test/jest-unit.json&quot;,\n    &quot;test:integration&quot;: &quot;jest --config ./test/jest-integration.json&quot;, \n    &quot;test:e2e&quot;: &quot;jest --config ./test/jest-e2e.json&quot;,\n    &quot;test:watch&quot;: &quot;jest --watch&quot;,\n    &quot;test:cov&quot;: &quot;jest --coverage&quot;,\n    \n    // 工具驱动的测试\n    &quot;test:api&quot;: &quot;newman run ./test/api-tests.postman_collection.json&quot;,\n    &quot;test:performance&quot;: &quot;artillery run ./test/load-test.yml&quot;,\n    &quot;test:security&quot;: &quot;npm audit &amp;&amp; snyk test&quot;,\n    &quot;test:lint&quot;: &quot;eslint src/**/*.ts&quot;,\n    \n    // 综合测试脚本\n    &quot;test:all&quot;: &quot;npm run test:unit &amp;&amp; npm run test:integration &amp;&amp; npm run test:e2e&quot;,\n    &quot;test:ci&quot;: &quot;npm run test:lint &amp;&amp; npm run test:security &amp;&amp; npm run test:all&quot;\n  }\n}\n</code></pre>\n<h3>6.4 测试策略的完整图景</h3>\n<pre><code class=\"language-plaintext\">代码驱动的自动化测试（开发者日常）    其他测试类型（专项/阶段性）\n                                   \n    E2E Tests (10%)                Manual Testing\n   ┌─────────────────┐              ┌─────────────────┐\n   │  关键业务流程    │              │  可用性、探索性  │\n   └─────────────────┘              └─────────────────┘\n                                   \n  Integration Tests (20%)          Tool-based Testing  \n ┌─────────────────────┐            ┌─────────────────┐\n │  模块间交互验证      │            │ 性能、安全扫描   │\n └─────────────────────┘            └─────────────────┘\n                                   \nUnit Tests (70%)                   Static Analysis\n┌─────────────────────┐             ┌─────────────────┐\n│  业务逻辑验证        │             │ 代码质量检查     │\n└─────────────────────┘             └─────────────────┘\n</code></pre>\n<h3>6.5 为什么本文重点讲代码驱动的测试</h3>\n<p><strong>开发者日常最需要的技能：</strong></p>\n<ol>\n<li>\n<p><strong>高频使用</strong>：每天开发过程中都要编写和运行</p>\n</li>\n<li>\n<p><strong>即时反馈</strong>：能在编码时立即发现问题</p>\n</li>\n<li>\n<p><strong>CI/CD集成</strong>：可以自动化集成到部署流程</p>\n</li>\n<li>\n<p><strong>成本效益</strong>：一次编写，持续受益</p>\n</li>\n</ol>\n<p><strong>其他测试类型的特点：</strong></p>\n<ul>\n<li>\n<p><strong>执行频率较低</strong>：通常在特定阶段执行（如发布前）</p>\n</li>\n<li>\n<p><strong>专门工具</strong>：需要学习和配置专门的测试工具</p>\n</li>\n<li>\n<p><strong>专业团队</strong>：更多由QA或运维团队负责</p>\n</li>\n<li>\n<p><strong>环境要求</strong>：需要特殊的测试环境和数据</p>\n</li>\n</ul>\n<h3>6.6 实际项目中的应用建议</h3>\n<p><strong>开发阶段（每日）：</strong></p>\n<ul>\n<li>\n<p>单元测试：验证业务逻辑</p>\n</li>\n<li>\n<p>集成测试：验证模块交互</p>\n</li>\n<li>\n<p>代码质量检查：ESLint、Prettier</p>\n</li>\n</ul>\n<p><strong>集成阶段（每次提交）：</strong></p>\n<ul>\n<li>\n<p>E2E测试：验证关键流程</p>\n</li>\n<li>\n<p>API测试：验证接口契约</p>\n</li>\n<li>\n<p>安全扫描：检查依赖漏洞</p>\n</li>\n</ul>\n<p><strong>发布阶段（版本发布前）：</strong></p>\n<ul>\n<li>\n<p>性能测试：验证系统负载能力</p>\n</li>\n<li>\n<p>兼容性测试：多浏览器/设备验证</p>\n</li>\n<li>\n<p>手工测试：用户体验验证</p>\n</li>\n</ul>\n<p>通过这种分层的测试策略，既保证了开发效率，又确保了产品质量。代码驱动的自动化测试构成了质量保障的基础，而其他测试类型则在特定场景下提供补充验证。</p>\n","date_published":"2025-07-20T00:00:00.000Z","tags":["Node","自动化测试","Nestjs测试","单元测试"],"language":"zh"},{"id":"https://www.lihuanyu.com/posts/2025/GitHub-Actions-%E8%87%AA%E5%8A%A8%E5%8F%91%E5%B8%83-npm-%E5%8C%85%E7%AE%80%E6%98%93%E6%8C%87%E5%8D%97/","url":"https://www.lihuanyu.com/posts/2025/GitHub-Actions-%E8%87%AA%E5%8A%A8%E5%8F%91%E5%B8%83-npm-%E5%8C%85%E7%AE%80%E6%98%93%E6%8C%87%E5%8D%97/","title":"GitHub Actions 自动发布 npm 包简易指南","summary":"已并入《GitHub Actions 适合做什么，不适合做什么》。","content_html":"<p>关于 GitHub Actions 自动发布 npm 包的内容，已经整理进更完整的文章：</p>\n<p><a href=\"/posts/2020/%E4%BB%8ETravis%E8%BF%81%E7%A7%BB%E5%88%B0GitHub-Actions/\">GitHub Actions 适合做什么，不适合做什么</a></p>\n<p>这页保留原链接，是因为 npm 包发布仍然是 GitHub Actions 非常适合的场景：它可以把 tag、测试、构建、发布和日志串成一个稳定流程。</p>\n<p>今天再配置 npm 发布时，除了传统的 <code>NPM_TOKEN</code>，也应该优先了解 npm trusted publishing。它通过 OIDC 建立 GitHub Actions 和 npm 之间的信任关系，减少长期 token 的暴露面。完整配置和适用条件应以 npm 官方文档为准。</p>\n","date_published":"2025-07-19T00:00:00.000Z","date_modified":"2026-05-04T00:00:00.000Z","tags":["Github Action","自动化","前端"],"language":"zh"},{"id":"https://www.lihuanyu.com/posts/2025/%E6%97%A0%E4%BA%BA%E9%A9%BE%E9%A9%B6%E4%B8%8E%E4%BA%BA%E5%9B%A0%E5%B7%A5%E7%A8%8B/","url":"https://www.lihuanyu.com/posts/2025/%E6%97%A0%E4%BA%BA%E9%A9%BE%E9%A9%B6%E4%B8%8E%E4%BA%BA%E5%9B%A0%E5%B7%A5%E7%A8%8B/","title":"无人驾驶与人因工程","summary":"从小米 SU7 高速事故谈起，借助情景意识和自动化接管问题讨论智能驾驶中的人因工程风险。","content_html":"<p>2025年3月29日22时44分，一辆小米SU7标准版在德上高速公路池祁段行驶过程中遭遇严重交通事故。根据小米公司披露的信息，事故发生前车辆处于NOA智能辅助驾驶状态，以116km/h时速持续行驶。事发路段因施工修缮，用路障封闭自车道、改道至逆向车道。车辆检测出障碍物后发出提醒并开始减速。随后驾驶员接管车辆进入人驾状态，持续减速并操控车辆转向，随后车辆与隔离带水泥桩发生碰撞，碰撞前系统最后可以确认的时速约为97km/h。</p>\n<p>小米汽车本身具有不小的话题性，事故出现后引起很多用户的争论，但很多人要么站队小米，指责驾驶员和其他批判小米的人，认为这起事故完全和小米无关，要么批判小米的技术，认为小米的技术不够好，不能保证安全。</p>\n<p>这种口水仗毫无意义，不如跳出这些话题，谈谈人因工程。</p>\n<p>首先尝试还原一下事故核心场景：</p>\n<p>随着筒锥的逼近，自动驾驶终于识别到前方有障碍物，告警提示驾驶员接管，但距离已经太近，驾驶员情急之下手打方向盘22度然后回正，最后与侧面护栏相撞。</p>\n<p><img src=\"https://aipaint.lihuanyu.com/2025-04-26/%E5%B0%8F%E7%B1%B3%E4%BA%8B%E6%95%85%E6%A8%A1%E6%8B%9F%E5%9B%BE.png\" alt=\"事故模拟还原\"></p>\n<p>整个轨迹大概如图：</p>\n<p><img src=\"https://aipaint.lihuanyu.com/2025-04-26/%E4%BA%8B%E6%95%85%E8%BD%A8%E8%BF%B9%E6%A8%A1%E6%8B%9F%E8%BF%98%E5%8E%9F.png\" alt=\"事故轨迹模拟\"></p>\n<p>如上图，事故其实是驾驶员为了避障过度转向导致的事故。</p>\n<p>但这并非驾驶员的错误。</p>\n<p>根据小米披露的数据，事发前驾驶员转向22度，刹车踏板开合角度31度。</p>\n<p><img src=\"https://aipaint.lihuanyu.com/2025-04-26/%E4%BA%8B%E6%95%85%E4%B8%AD%E6%96%B9%E5%90%91%E7%9B%98%E4%B8%8E%E5%88%B9%E8%BD%A6%E6%95%B0%E6%8D%AE.png\" alt=\"方向盘与刹车数据\"></p>\n<p>在高速路上，方向盘打22度是什么概念？一般正常变道或轻微转向通常只需要约5-15度的方向盘转动角度。紧急避险情况下，方向盘转动角度可能会达到20-30度，但这已经是相当大的转向幅度，会导致车辆明显的横向移动。高速行驶时（如100km/h以上）方向盘转动超过30度就已经是非常剧烈的转向了，可能导致车辆失控。</p>\n<p>也就是说，这个转动幅度很大，但并非完全不能操作。那看起来好像还是驾驶员的经验与能力的问题？</p>\n<p>然而，将责任归咎于驾驶员是不公平的。这涉及到人因工程中的一个核心概念：情景意识（Situation Awareness）。在这种紧急情况下，我们不能期望驾驶员能够立即建立完整的情景意识。</p>\n<p>传统手动驾驶中，驾驶员需要不断进行微调方向盘，这个过程帮助大脑建立了速度与转向角度之间的对应关系。这种持续的反馈形成了驾驶员的情景意识，使他们能够准确判断在特定速度下需要多大的转向角度。</p>\n<p>智能辅助驾驶系统虽然减轻了驾驶员的负担，但同时也切断了这种持续反馈。当系统突然要求驾驶员接管时，驾驶员缺乏当前情境下的&quot;感觉&quot;，只能依靠长期记忆中的经验来判断。</p>\n<p>在非紧急情况下，驾驶员通常会先尝试小角度转向，然后根据车辆反应逐渐调整。但当障碍物近在眼前时，驾驶员必须立即给出一个&quot;足够大&quot;的转向角度，<strong>而这个判断往往不够准确</strong>。本次事故中，驾驶员给出的22度转向角度就是这种紧急情况下的本能反应。</p>\n<p>这种现象在航空领域有着更为惨痛的教训。2009年的法航447航班空难就是一个典型案例：当空速管结冰导致自动驾驶突然退出时，接手的飞行员缺乏对当前飞行状态的准确感知，做出了错误的操作决策，最终导致飞机坠毁，228人全部遇难。</p>\n<p>这类事故之所以反复发生，是因为现代自动化系统往往将人排除在控制回路之外。系统正常运行时，操作员不需要（也不被要求）了解系统的所有细节。随着时间推移，操作员对系统状态的理解逐渐过时，当系统突然要求人工接管时，操作员需要时间重新建立情景意识，而紧急情况往往不给这个时间。</p>\n<p>人因工程学界将这种现象称为&quot;伐木工效应&quot;（Lumberjack Effect）：自动化程度越高，操作员在日常中的参与度就越低，技能保持越差，当需要接管时就越容易出错。这是一个悖论：自动化系统越先进，在罕见的需要人工接管的情况下，失败的风险反而更高。</p>\n<p>解决这一问题的方向包括：设计更透明的自动化系统，让操作员始终了解系统状态；开发自适应自动化，根据情况动态调整自动化水平；将自动化系统设计为&quot;团队成员&quot;而非替代者，保持人在回路中的参与。</p>\n<blockquote>\n<p>让我想起了《流浪地球2》里最后的彩蛋，提到moss的训练模式引入了 人在回路 。</p>\n</blockquote>\n<p>遗憾的是，工业界特别是国内对人因工程的重视程度不够。人因工程常被视为&quot;软科学&quot;，甚至被贬低为&quot;文科&quot;。当系统设计要求人类做出超出人类能力范围的操作时，事故责任往往被归咎于操作员&quot;不够专注&quot;或&quot;训练不足&quot;。</p>\n<p>只有航空、核电等对安全有极高要求的行业才真正重视人因工程。对于新兴的智能驾驶领域，这方面的进步可能还需要付出更多代价才能实现。</p>\n","date_published":"2025-04-26T00:00:00.000Z","tags":["随笔","人因工程","无人驾驶"],"language":"zh"},{"id":"https://www.lihuanyu.com/posts/2025/%E5%A6%82%E4%BD%95%E9%9B%86%E6%88%90Github%E7%99%BB%E5%BD%95/","url":"https://www.lihuanyu.com/posts/2025/%E5%A6%82%E4%BD%95%E9%9B%86%E6%88%90Github%E7%99%BB%E5%BD%95/","title":"GitHub OAuth 登录教程：PKCE、state 与 Session","summary":"使用 GitHub OAuth Authorization Code flow 接入登录，包含 state 校验、PKCE、服务端换取 token、用户绑定和本站 Session。","content_html":"<p>GitHub 登录应由服务端完成授权码换 token、用户身份查询和本站 Session 创建。浏览器只负责跳转与接收 HttpOnly Session Cookie，不应持有 <code>client_secret</code> 或 GitHub access token。</p>\n<p>本文使用 GitHub OAuth App 的 Web application flow，并同时启用 <code>state</code> 与 Proof Key for Code Exchange（PKCE）。这套结构适合开发者工具、技术社区和开源项目后台。</p>\n<p><a href=\"/en/posts/2025/implement-login-with-github-safely/\">English version: How to Implement Login with GitHub Safely</a></p>\n<p>接入完成后，各层只保留自己需要的数据：</p>\n<div class=\"table-scroll\"><table>\n<thead>\n<tr>\n<th>位置</th>\n<th>保存内容</th>\n<th>不应保存</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>浏览器</td>\n<td>本站 HttpOnly Session Cookie</td>\n<td><code>client_secret</code>、GitHub access token</td>\n</tr>\n<tr>\n<td>服务端临时会话</td>\n<td><code>state</code>、<code>code_verifier</code></td>\n<td>长期有效的明文登录凭据</td>\n</tr>\n<tr>\n<td>服务端用户表</td>\n<td>GitHub user id、本站用户关系</td>\n<td>把可修改的 GitHub login 当唯一身份</td>\n</tr>\n</tbody>\n</table>\n</div><h2>先确定 OAuth 登录边界</h2>\n<p>GitHub OAuth App 的 Web application flow 大致是三步：</p>\n<ol>\n<li>用户从你的站点跳转到 GitHub 授权页。</li>\n<li>GitHub 授权后带着临时 <code>code</code> 和 <code>state</code> 跳回你的回调地址。</li>\n<li>服务端用 <code>code</code> 换取 access token，再用 token 请求 GitHub API 获取用户身份。</li>\n</ol>\n<p>这里最关键的是两点：</p>\n<ul>\n<li><code>client_secret</code> 只能放在服务端。</li>\n<li><code>state</code> 必须校验，用来防止 CSRF 和错误会话串联。</li>\n</ul>\n<p>如果只是做“用 GitHub 身份登录本站”，OAuth App 已经够用。如果需要更细粒度的仓库权限、安装到组织、以应用身份执行自动化任务，就应该优先评估 GitHub App。</p>\n<h2>创建 GitHub OAuth App</h2>\n<p>在 GitHub 里进入：</p>\n<pre><code class=\"language-text\">Settings -&gt; Developer settings -&gt; OAuth Apps -&gt; New OAuth App\n</code></pre>\n<p>需要填写几个关键字段：</p>\n<ul>\n<li><code>Application name</code>：应用名称。</li>\n<li><code>Homepage URL</code>：站点首页地址。</li>\n<li><code>Authorization callback URL</code>：授权回调地址，比如 <code>https://example.com/auth/github/callback</code>。</li>\n</ul>\n<p>创建完成后会得到：</p>\n<ul>\n<li><code>Client ID</code>：可以出现在授权 URL 里。</li>\n<li><code>Client Secret</code>：必须只保存在服务端，通常放在环境变量或密钥管理系统里。</li>\n</ul>\n<p>本地开发时可以把 callback 配成：</p>\n<pre><code class=\"language-text\">http://localhost:3000/auth/github/callback\n</code></pre>\n<p>线上环境要使用 HTTPS。</p>\n<h2>使用 Authorization Code 和 PKCE 流程</h2>\n<p>一个比较清晰的登录链路是：</p>\n<pre><code class=\"language-text\">浏览器点击 GitHub 登录\n  -&gt; 服务端生成 state 和 code_verifier\n  -&gt; 服务端把 state/code_verifier 写入 HttpOnly 临时 cookie 或 session\n  -&gt; 服务端重定向到 GitHub 授权页\n  -&gt; GitHub 回调 /auth/github/callback?code=...&amp;state=...\n  -&gt; 服务端校验 state\n  -&gt; 服务端用 code 换 access token\n  -&gt; 服务端请求 GitHub /user 和 /user/emails\n  -&gt; 服务端创建或更新本地用户\n  -&gt; 服务端写入本站 session cookie\n  -&gt; 浏览器回到业务页面\n</code></pre>\n<p>GitHub 文档把 <code>code_challenge</code> 和 <code>code_verifier</code> 标为强烈推荐，并要求 challenge method 使用 <code>S256</code>。即使服务端应用已经有 <code>client_secret</code>，PKCE 仍能降低授权码被截获后的风险。</p>\n<h2>从服务端发起授权请求</h2>\n<p>下面用 Express 写一个示例。真实项目里可以把临时数据放进 Redis、数据库 session 或加密 cookie。</p>\n<pre><code class=\"language-js\">import crypto from 'node:crypto';\nimport express from 'express';\nimport cookieParser from 'cookie-parser';\n\nconst app = express();\napp.use(cookieParser());\n\nconst clientId = process.env.GITHUB_CLIENT_ID;\nconst clientSecret = process.env.GITHUB_CLIENT_SECRET;\nconst redirectUri = 'http://localhost:3000/auth/github/callback';\nconst isProduction = process.env.NODE_ENV === 'production';\n\nfunction base64url(buffer) {\n  return buffer\n    .toString('base64')\n    .replace(/\\+/g, '-')\n    .replace(/\\//g, '_')\n    .replace(/=+$/g, '');\n}\n\nfunction createCodeChallenge(verifier) {\n  return base64url(crypto.createHash('sha256').update(verifier).digest());\n}\n\napp.get('/auth/github/start', (req, res) =&gt; {\n  const state = base64url(crypto.randomBytes(32));\n  const codeVerifier = base64url(crypto.randomBytes(32));\n  const codeChallenge = createCodeChallenge(codeVerifier);\n\n  res.cookie('github_oauth_state', state, {\n    httpOnly: true,\n    secure: isProduction,\n    sameSite: 'lax',\n    maxAge: 10 * 60 * 1000,\n  });\n\n  res.cookie('github_oauth_code_verifier', codeVerifier, {\n    httpOnly: true,\n    secure: isProduction,\n    sameSite: 'lax',\n    maxAge: 10 * 60 * 1000,\n  });\n\n  const params = new URLSearchParams({\n    client_id: clientId,\n    redirect_uri: redirectUri,\n    scope: 'read:user user:email',\n    state,\n    code_challenge: codeChallenge,\n    code_challenge_method: 'S256',\n  });\n\n  res.redirect(`https://github.com/login/oauth/authorize?${params}`);\n});\n</code></pre>\n<p><code>scope</code> 不要贪多。只需要登录身份时，常见选择是：</p>\n<ul>\n<li><code>read:user</code>：读取基础用户资料。</li>\n<li><code>user:email</code>：读取用户邮箱，尤其是主资料里的 <code>email</code> 为空时。</li>\n</ul>\n<p>不要为了登录直接申请 <code>repo</code> 这类高权限 scope。权限越大，用户越警惕，token 泄露后的风险也越大。</p>\n<h2>校验回调并创建本站 Session</h2>\n<p>GitHub 回调时会带上 <code>code</code> 和 <code>state</code>。服务端必须先检查 <code>state</code> 是否和自己之前保存的一致，不一致就终止流程。</p>\n<pre><code class=\"language-js\">app.get('/auth/github/callback', async (req, res) =&gt; {\n  const { code, state, error } = req.query;\n\n  if (error || !code || !state) {\n    return res.status(400).send('GitHub authorization was cancelled or invalid');\n  }\n\n  const expectedState = req.cookies.github_oauth_state;\n  const codeVerifier = req.cookies.github_oauth_code_verifier;\n\n  res.clearCookie('github_oauth_state');\n  res.clearCookie('github_oauth_code_verifier');\n\n  if (state !== expectedState || !codeVerifier) {\n    return res.status(400).send('Invalid OAuth state');\n  }\n\n  const tokenResponse = await fetch('https://github.com/login/oauth/access_token', {\n    method: 'POST',\n    headers: {\n      Accept: 'application/json',\n      'Content-Type': 'application/x-www-form-urlencoded',\n    },\n    body: new URLSearchParams({\n      client_id: clientId,\n      client_secret: clientSecret,\n      code: String(code),\n      redirect_uri: redirectUri,\n      code_verifier: codeVerifier,\n    }),\n  });\n\n  const tokenData = await tokenResponse.json();\n\n  if (!tokenResponse.ok || tokenData.error) {\n    return res.status(401).json({\n      message: 'GitHub authorization failed',\n      error: tokenData.error,\n    });\n  }\n\n  const accessToken = tokenData.access_token;\n\n  const githubUser = await fetchGitHubUser(accessToken);\n  const emails = await fetchGitHubEmails(accessToken);\n\n  const primaryEmail =\n    emails.find((email) =&gt; email.primary &amp;&amp; email.verified)?.email ??\n    githubUser.email;\n\n  const user = await upsertUserFromGitHub({\n    githubId: githubUser.id,\n    login: githubUser.login,\n    name: githubUser.name,\n    avatarUrl: githubUser.avatar_url,\n    email: primaryEmail,\n  });\n\n  const sessionId = await createSession(user.id);\n\n  res.cookie('session_id', sessionId, {\n    httpOnly: true,\n    secure: isProduction,\n    sameSite: 'lax',\n  });\n\n  res.redirect('/dashboard');\n});\n</code></pre>\n<p>请求 GitHub API 时使用 <code>Authorization: Bearer</code>：</p>\n<pre><code class=\"language-js\">async function fetchGitHubUser(accessToken) {\n  const response = await fetch('https://api.github.com/user', {\n    headers: {\n      Authorization: `Bearer ${accessToken}`,\n      Accept: 'application/vnd.github+json',\n      'X-GitHub-Api-Version': '2022-11-28',\n    },\n  });\n\n  if (!response.ok) {\n    throw new Error('Failed to fetch GitHub user');\n  }\n\n  return response.json();\n}\n\nasync function fetchGitHubEmails(accessToken) {\n  const response = await fetch('https://api.github.com/user/emails', {\n    headers: {\n      Authorization: `Bearer ${accessToken}`,\n      Accept: 'application/vnd.github+json',\n      'X-GitHub-Api-Version': '2022-11-28',\n    },\n  });\n\n  if (!response.ok) {\n    return [];\n  }\n\n  return response.json();\n}\n</code></pre>\n<p><code>upsertUserFromGitHub()</code> 和 <code>createSession()</code> 取决于自己的业务系统。常见做法是：</p>\n<ul>\n<li>用 GitHub user id 绑定本地用户，而不是只用 login。login 可能改名，id 更稳定。</li>\n<li>保存头像、昵称、邮箱等展示字段。</li>\n<li>用自己的 session 或 JWT 管理本站登录态。</li>\n<li>GitHub token 如果后续不需要调用 GitHub API，就不要长期保存。</li>\n</ul>\n<h2>让前端只负责跳转</h2>\n<p>前端只需要把用户带到服务端的登录入口：</p>\n<pre><code class=\"language-html\">&lt;a href=&quot;/auth/github/start&quot;&gt;Continue with GitHub&lt;/a&gt;\n</code></pre>\n<p>或者按钮点击后跳转：</p>\n<pre><code class=\"language-js\">document.querySelector('#github-login').addEventListener('click', () =&gt; {\n  window.location.href = '/auth/github/start';\n});\n</code></pre>\n<p>前端不需要知道 <code>client_secret</code>，也不应该把 GitHub access token 存进 <code>localStorage</code>。浏览器侧只持有本站自己的登录态 cookie。</p>\n<h2>避免常见 OAuth 登录错误</h2>\n<h3>没有校验 state</h3>\n<p><code>state</code> 是 OAuth 登录里最容易被省略、也最不该省略的字段。它应该是不可猜测的随机字符串，并且和当前登录发起方绑定。回调时如果不一致，流程必须终止。</p>\n<h3>把 token 返回给前端</h3>\n<p>GitHub access token 代表用户授权。把它返回给前端并存入 <code>localStorage</code>，会扩大 XSS 后的损失。除非是纯前端应用且做了专门设计，否则更推荐服务端持有 token，并给浏览器发本站 session。</p>\n<h3>用 login 当唯一身份</h3>\n<p>GitHub 用户名可以修改。数据库绑定用户时应该优先使用 GitHub user id。</p>\n<h3>scope 申请过大</h3>\n<p>登录通常不需要仓库权限。权限申请越大，授权页面越吓人，也越难通过用户信任。</p>\n<h3>忽略邮箱为空</h3>\n<p>GitHub 用户资料里的 <code>email</code> 可能为空。需要邮箱时，要通过 <code>user:email</code> scope 调 <code>/user/emails</code>，并优先选择已验证的主邮箱。</p>\n<h2>把 GitHub 当身份来源而不是本站会话</h2>\n<p>GitHub 登录的核心不是在前端拼一个授权 URL，而是把 OAuth 的安全边界放对：</p>\n<ul>\n<li>前端负责跳转。</li>\n<li>服务端保存 <code>client_secret</code>。</li>\n<li><code>state</code> 用来绑定登录请求和回调。</li>\n<li>授权码在服务端换 token。</li>\n<li>token 用来向 GitHub 确认用户身份。</li>\n<li>本站登录态由自己的 session 系统管理。</li>\n</ul>\n<p>这样接入后，GitHub 只是身份提供方，真正的账号体系仍然掌握在自己的应用里。</p>\n<p>如果多个产品需要共享登录、统一账号和跨应用 Session，问题会从“接入一个 OAuth 提供方”升级为身份中心设计，可以继续阅读 <a href=\"/posts/oidc-identity-boundary-behind-login/\">OIDC：登录背后的身份边界</a>。</p>\n<h2>扩展阅读</h2>\n<ul>\n<li><a href=\"https://docs.github.com/en/apps/oauth-apps/building-oauth-apps/authorizing-oauth-apps\">GitHub Docs: Authorizing OAuth apps</a></li>\n<li><a href=\"https://docs.github.com/en/apps/oauth-apps/building-oauth-apps/creating-an-oauth-app\">GitHub Docs: Creating an OAuth app</a></li>\n<li><a href=\"https://docs.github.com/en/apps/oauth-apps/building-oauth-apps/scopes-for-oauth-apps\">GitHub Docs: Scopes for OAuth apps</a></li>\n</ul>\n","date_published":"2025-04-20T00:00:00.000Z","date_modified":"2026-08-01T00:00:00.000Z","tags":["前端","OAuth","Github","登录","OAuth2.0","Github登录","三方登录"],"language":"zh"},{"id":"https://www.lihuanyu.com/en/posts/2025/implement-login-with-github-safely/","url":"https://www.lihuanyu.com/en/posts/2025/implement-login-with-github-safely/","title":"How to Implement Login with GitHub Safely","summary":"A practical server-side GitHub OAuth login flow using state validation, authorization code exchange, GitHub user lookup, and a local session.","content_html":"<p>Many websites no longer build a complete username and password system from scratch. They use third-party login instead. For developer tools, technical communities, and open source project dashboards, GitHub login is a common choice.</p>\n<p>But “implement Login with GitHub” is sometimes misunderstood as “let the frontend obtain a GitHub access token and store it in the browser.” That can make a demo work, but it is not a good default for a real application.</p>\n<p>A safer structure is: the browser handles redirects, while the server validates <code>state</code>, exchanges the authorization code for a token, fetches the GitHub user identity, and creates the application’s own login session. The frontend receives your site’s session, not the GitHub token.</p>\n<p><a href=\"/posts/2025/%E5%A6%82%E4%BD%95%E9%9B%86%E6%88%90Github%E7%99%BB%E5%BD%95/\">Chinese version of this article</a></p>\n<h2>OAuth Flow Responsibilities</h2>\n<p>GitHub OAuth App’s web application flow has three main steps:</p>\n<ol>\n<li>The user is redirected from your site to GitHub’s authorization page.</li>\n<li>After authorization, GitHub redirects back to your callback URL with a temporary <code>code</code> and the <code>state</code> value.</li>\n<li>Your server exchanges the <code>code</code> for an access token, then uses that token to call the GitHub API and identify the user.</li>\n</ol>\n<p>Two details are critical:</p>\n<ul>\n<li><code>client_secret</code> belongs only on the server.</li>\n<li><code>state</code> must be validated to protect against CSRF and mixed-up login sessions.</li>\n</ul>\n<p>If the only goal is to let users sign in with their GitHub identity, an OAuth App is usually enough. If you need fine-grained repository permissions, organization installation, or automation as an app identity, evaluate GitHub Apps first.</p>\n<h2>Create a GitHub OAuth App</h2>\n<p>In GitHub, open:</p>\n<pre><code class=\"language-text\">Settings -&gt; Developer settings -&gt; OAuth Apps -&gt; New OAuth App\n</code></pre>\n<p>Fill in the key fields:</p>\n<ul>\n<li><code>Application name</code>: the app name.</li>\n<li><code>Homepage URL</code>: your site homepage.</li>\n<li><code>Authorization callback URL</code>: for example, <code>https://example.com/auth/github/callback</code>.</li>\n</ul>\n<p>After creation, GitHub gives you:</p>\n<ul>\n<li><code>Client ID</code>: safe to include in the authorization URL.</li>\n<li><code>Client Secret</code>: server-only, usually stored in environment variables or a secret manager.</li>\n</ul>\n<p>For local development, the callback URL can be:</p>\n<pre><code class=\"language-text\">http://localhost:3000/auth/github/callback\n</code></pre>\n<p>Use HTTPS in production.</p>\n<h2>Recommended Architecture</h2>\n<p>A clean login flow looks like this:</p>\n<pre><code class=\"language-text\">Browser clicks Login with GitHub\n  -&gt; server generates state and code_verifier\n  -&gt; server stores state/code_verifier in an HttpOnly temporary cookie or session\n  -&gt; server redirects to GitHub authorization page\n  -&gt; GitHub redirects to /auth/github/callback?code=...&amp;state=...\n  -&gt; server validates state\n  -&gt; server exchanges code for access token\n  -&gt; server requests GitHub /user and /user/emails\n  -&gt; server creates or updates local user\n  -&gt; server writes local session cookie\n  -&gt; browser returns to the application page\n</code></pre>\n<p>PKCE can be used here too. GitHub’s documentation now strongly recommends <code>code_challenge</code> and <code>code_verifier</code>. Even when a server-side application already has a <code>client_secret</code>, PKCE still reduces the risk if an authorization code is intercepted.</p>\n<h2>Start the Authorization Request</h2>\n<p>The example below uses Express. In a real project, temporary OAuth state can live in Redis, a database session, or an encrypted cookie.</p>\n<pre><code class=\"language-js\">import crypto from 'node:crypto';\nimport express from 'express';\nimport cookieParser from 'cookie-parser';\n\nconst app = express();\napp.use(cookieParser());\n\nconst clientId = process.env.GITHUB_CLIENT_ID;\nconst clientSecret = process.env.GITHUB_CLIENT_SECRET;\nconst redirectUri = 'http://localhost:3000/auth/github/callback';\nconst isProduction = process.env.NODE_ENV === 'production';\n\nfunction base64url(buffer) {\n  return buffer\n    .toString('base64')\n    .replace(/\\+/g, '-')\n    .replace(/\\//g, '_')\n    .replace(/=+$/g, '');\n}\n\nfunction createCodeChallenge(verifier) {\n  return base64url(crypto.createHash('sha256').update(verifier).digest());\n}\n\napp.get('/auth/github/start', (req, res) =&gt; {\n  const state = base64url(crypto.randomBytes(32));\n  const codeVerifier = base64url(crypto.randomBytes(32));\n  const codeChallenge = createCodeChallenge(codeVerifier);\n\n  res.cookie('github_oauth_state', state, {\n    httpOnly: true,\n    secure: isProduction,\n    sameSite: 'lax',\n    maxAge: 10 * 60 * 1000,\n  });\n\n  res.cookie('github_oauth_code_verifier', codeVerifier, {\n    httpOnly: true,\n    secure: isProduction,\n    sameSite: 'lax',\n    maxAge: 10 * 60 * 1000,\n  });\n\n  const params = new URLSearchParams({\n    client_id: clientId,\n    redirect_uri: redirectUri,\n    scope: 'read:user user:email',\n    state,\n    code_challenge: codeChallenge,\n    code_challenge_method: 'S256',\n  });\n\n  res.redirect(`https://github.com/login/oauth/authorize?${params}`);\n});\n</code></pre>\n<p>Keep scopes small. For login identity, common scopes are:</p>\n<ul>\n<li><code>read:user</code>: read basic user profile data.</li>\n<li><code>user:email</code>: read user email addresses, especially when the profile-level <code>email</code> field is empty.</li>\n</ul>\n<p>Do not request high-permission scopes such as <code>repo</code> just for login. Larger scopes make users more cautious and increase the damage if a token leaks.</p>\n<h2>Handle the GitHub Callback</h2>\n<p>GitHub redirects back with <code>code</code> and <code>state</code>. The server must first compare the returned <code>state</code> with the value it stored earlier. If they do not match, abort the flow.</p>\n<pre><code class=\"language-js\">app.get('/auth/github/callback', async (req, res) =&gt; {\n  const { code, state } = req.query;\n\n  if (!code || !state) {\n    return res.status(400).send('Missing OAuth code or state');\n  }\n\n  if (state !== req.cookies.github_oauth_state) {\n    return res.status(400).send('Invalid OAuth state');\n  }\n\n  const tokenResponse = await fetch('https://github.com/login/oauth/access_token', {\n    method: 'POST',\n    headers: {\n      Accept: 'application/json',\n      'Content-Type': 'application/x-www-form-urlencoded',\n    },\n    body: new URLSearchParams({\n      client_id: clientId,\n      client_secret: clientSecret,\n      code: String(code),\n      redirect_uri: redirectUri,\n      code_verifier: req.cookies.github_oauth_code_verifier,\n    }),\n  });\n\n  const tokenData = await tokenResponse.json();\n\n  if (!tokenResponse.ok || tokenData.error) {\n    return res.status(401).json({\n      message: 'GitHub authorization failed',\n      error: tokenData.error,\n    });\n  }\n\n  const accessToken = tokenData.access_token;\n\n  const githubUser = await fetchGitHubUser(accessToken);\n  const emails = await fetchGitHubEmails(accessToken);\n\n  const primaryEmail =\n    emails.find((email) =&gt; email.primary &amp;&amp; email.verified)?.email ??\n    githubUser.email;\n\n  const user = await upsertUserFromGitHub({\n    githubId: githubUser.id,\n    login: githubUser.login,\n    name: githubUser.name,\n    avatarUrl: githubUser.avatar_url,\n    email: primaryEmail,\n  });\n\n  const sessionId = await createSession(user.id);\n\n  res.clearCookie('github_oauth_state');\n  res.clearCookie('github_oauth_code_verifier');\n  res.cookie('session_id', sessionId, {\n    httpOnly: true,\n    secure: isProduction,\n    sameSite: 'lax',\n  });\n\n  res.redirect('/dashboard');\n});\n</code></pre>\n<p>Use <code>Authorization: Bearer</code> when calling the GitHub API:</p>\n<pre><code class=\"language-js\">async function fetchGitHubUser(accessToken) {\n  const response = await fetch('https://api.github.com/user', {\n    headers: {\n      Authorization: `Bearer ${accessToken}`,\n      Accept: 'application/vnd.github+json',\n    },\n  });\n\n  if (!response.ok) {\n    throw new Error('Failed to fetch GitHub user');\n  }\n\n  return response.json();\n}\n\nasync function fetchGitHubEmails(accessToken) {\n  const response = await fetch('https://api.github.com/user/emails', {\n    headers: {\n      Authorization: `Bearer ${accessToken}`,\n      Accept: 'application/vnd.github+json',\n    },\n  });\n\n  if (!response.ok) {\n    return [];\n  }\n\n  return response.json();\n}\n</code></pre>\n<p><code>upsertUserFromGitHub()</code> and <code>createSession()</code> depend on your own application. Common practices are:</p>\n<ul>\n<li>Bind local users by GitHub user id, not only by login. A login can change; the id is more stable.</li>\n<li>Store display fields such as avatar, name, and email.</li>\n<li>Use your own session or JWT system for your site.</li>\n<li>Do not store the GitHub token long term if you do not need to call GitHub APIs later.</li>\n</ul>\n<h2>What the Frontend Should Do</h2>\n<p>The frontend only needs to send users to the server-side login entry:</p>\n<pre><code class=\"language-html\">&lt;a href=&quot;/auth/github/start&quot;&gt;Continue with GitHub&lt;/a&gt;\n</code></pre>\n<p>Or redirect on button click:</p>\n<pre><code class=\"language-js\">document.querySelector('#github-login').addEventListener('click', () =&gt; {\n  window.location.href = '/auth/github/start';\n});\n</code></pre>\n<p>The frontend does not need to know <code>client_secret</code>, and it should not store the GitHub access token in <code>localStorage</code>. The browser should hold only your application’s own session cookie.</p>\n<h2>Common Pitfalls</h2>\n<h3>Skipping state Validation</h3>\n<p><code>state</code> is easy to omit and should not be omitted. It should be an unguessable random string tied to the login attempt. If the callback value does not match, abort the flow.</p>\n<h3>Returning the Token to the Frontend</h3>\n<p>A GitHub access token represents user authorization. Returning it to the frontend and storing it in <code>localStorage</code> increases the damage of an XSS issue. Unless the application is intentionally designed as a pure frontend OAuth client, prefer server-held tokens and browser-held local sessions.</p>\n<h3>Using login as the Only Identifier</h3>\n<p>GitHub usernames can change. Prefer GitHub user id when binding accounts in your database.</p>\n<h3>Requesting Too Many Scopes</h3>\n<p>Login usually does not require repository access. Larger scopes make the authorization page look more sensitive and make user trust harder to earn.</p>\n<h3>Assuming email Is Always Present</h3>\n<p>The <code>email</code> field on the GitHub user profile can be empty. If your application needs email, request <code>user:email</code>, call <code>/user/emails</code>, and prefer a verified primary email.</p>\n<h2>Conclusion</h2>\n<p>The core of GitHub login is not building an authorization URL in the frontend. It is putting the OAuth security boundary in the right place:</p>\n<ul>\n<li>The frontend redirects.</li>\n<li>The server stores <code>client_secret</code>.</li>\n<li><code>state</code> binds the login request to the callback.</li>\n<li>The authorization code is exchanged on the server.</li>\n<li>The token is used to validate the user’s GitHub identity.</li>\n<li>Your own application session manages the logged-in state.</li>\n</ul>\n<p>With this structure, GitHub is the identity provider, while your application still owns its account system.</p>\n<h2>Further Reading</h2>\n<ul>\n<li><a href=\"https://docs.github.com/en/apps/oauth-apps/building-oauth-apps/authorizing-oauth-apps\">GitHub Docs: Authorizing OAuth apps</a></li>\n<li><a href=\"https://docs.github.com/en/apps/oauth-apps/building-oauth-apps/creating-an-oauth-app\">GitHub Docs: Creating an OAuth app</a></li>\n<li><a href=\"https://docs.github.com/en/apps/oauth-apps/building-oauth-apps/scopes-for-oauth-apps\">GitHub Docs: Scopes for OAuth apps</a></li>\n</ul>\n","date_published":"2025-04-20T00:00:00.000Z","date_modified":"2026-05-05T00:00:00.000Z","tags":["Frontend","OAuth","GitHub","Login","OAuth 2.0","Authentication"],"language":"en"},{"id":"https://www.lihuanyu.com/posts/2025/%E4%BA%92%E8%81%94%E7%BD%91%E5%88%9B%E4%B8%9A%E5%AF%92%E5%86%AC/","url":"https://www.lihuanyu.com/posts/2025/%E4%BA%92%E8%81%94%E7%BD%91%E5%88%9B%E4%B8%9A%E5%AF%92%E5%86%AC/","title":"互联网创业寒冬","summary":"已并入《平台、算法与创作者：为什么还需要独立博客》。","content_html":"<p>关于互联网创业寒冬、平台成熟、增长变难和个人创作者处境的思考，已经整理进更完整的文章：</p>\n<p><a href=\"/posts/2025/%E5%9C%A8%E5%9B%BD%E5%86%85%E7%9A%84%E5%B9%B3%E5%8F%B0%E4%BD%A0%E6%B2%A1%E6%9C%89%E7%B2%89%E4%B8%9D/\">平台、算法与创作者：为什么还需要独立博客</a></p>\n<p>这页保留原链接，是因为创业寒冬和创作者寒冬有相似的底层逻辑：早期增长红利消退后，产品和内容都不能再假设“只要足够好就会自然增长”。增长本身已经变成单独的问题，而平台掌握着关键入口。</p>\n<p><img src=\"https://aipaint.lihuanyu.com/2025-02-16/civitai.png\" alt=\"寒冬与巨鲸\"></p>\n<p>完整文章更关注个人层面的应对：继续使用平台获取曝光，但把长期内容、稳定 URL、上下文和可迁移资产沉淀到独立博客。</p>\n","date_published":"2025-02-16T00:00:00.000Z","date_modified":"2026-05-04T00:00:00.000Z","tags":["随笔","创业","互联网"],"language":"zh"},{"id":"https://www.lihuanyu.com/posts/2025/%E5%9C%A8Github%20Action%E9%87%8C%E6%9E%84%E5%BB%BA%E5%A4%A7%E5%9E%8BDocker%E9%95%9C%E5%83%8F/","url":"https://www.lihuanyu.com/posts/2025/%E5%9C%A8Github%20Action%E9%87%8C%E6%9E%84%E5%BB%BA%E5%A4%A7%E5%9E%8BDocker%E9%95%9C%E5%83%8F/","title":"在 GitHub Actions 里构建大型 Docker 镜像","summary":"已并入《GitHub Actions 适合做什么，不适合做什么》。","content_html":"<p>关于在 GitHub Actions 里构建 Docker 镜像，以及它和服务器部署边界的关系，已经整理进更完整的文章：</p>\n<p><a href=\"/posts/2020/%E4%BB%8ETravis%E8%BF%81%E7%A7%BB%E5%88%B0GitHub-Actions/\">GitHub Actions 适合做什么，不适合做什么</a></p>\n<p>这页保留原链接，是因为 Docker 镜像构建仍然是 GitHub Actions 很实用的场景。普通 Web 服务镜像适合在 Actions 里构建并推送到镜像仓库，让服务器只负责拉取和运行。</p>\n<p>但大型 AI 镜像不一样。Stable Diffusion、PyTorch、CUDA 等依赖会很快碰到 runner 磁盘和内存边界。清理 runner 空间可以解决一部分问题，但不是长期方案。镜像继续变大时，更应该考虑优化 Dockerfile、使用缓存、使用 larger runner、自托管 runner，或者把构建放到更靠近目标环境的专用机器上。</p>\n","date_published":"2025-02-16T00:00:00.000Z","date_modified":"2026-05-04T00:00:00.000Z","tags":["Github","Docker"],"language":"zh"},{"id":"https://www.lihuanyu.com/posts/2025/%E5%B0%8F%E7%A8%8B%E5%BA%8F%E7%BB%84%E4%BB%B6%E5%BA%93%E8%AF%A5%E7%94%A8rpx%E8%BF%98%E6%98%AFpx/","url":"https://www.lihuanyu.com/posts/2025/%E5%B0%8F%E7%A8%8B%E5%BA%8F%E7%BB%84%E4%BB%B6%E5%BA%93%E8%AF%A5%E7%94%A8rpx%E8%BF%98%E6%98%AFpx/","title":"小程序组件库用 px 还是 rpx？兼容性与选型指南","summary":"解释微信小程序 px 与 rpx 的换算规则、translateX 等场景的单位行为，并给出业务页面、自研组件库和第三方组件库的选型建议。","content_html":"<p>小程序组件库不需要在 <code>px</code> 和 <code>rpx</code> 之间二选一。更稳妥的默认方案是：用 <code>rpx</code> 表达随屏幕宽度变化的布局，用 <code>px</code> 表达不希望被大屏持续放大的尺寸，并把单位选择写进组件的设计约定。</p>\n<p>不要对第三方组件库做全量 <code>px -&gt; rpx</code> 转换。全量转换会同时改变边框、图标、弹层和内部定位，升级组件库时也更难判断差异。</p>\n<h2>先理解 px 和 rpx 的换算关系</h2>\n<p>微信小程序把屏幕宽度定义为 <code>750rpx</code>。换算关系是：</p>\n<pre><code class=\"language-text\">1rpx = windowWidth / 750 px\n</code></pre>\n<p>不同窗口宽度下，同一个 <code>rpx</code> 值对应不同的 CSS 像素：</p>\n<div class=\"table-scroll\"><table>\n<thead>\n<tr>\n<th style=\"text-align:right\">窗口宽度</th>\n<th style=\"text-align:right\">1rpx 对应的 px</th>\n<th style=\"text-align:right\">750rpx 的实际宽度</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td style=\"text-align:right\">320px</td>\n<td style=\"text-align:right\">0.427px</td>\n<td style=\"text-align:right\">320px</td>\n</tr>\n<tr>\n<td style=\"text-align:right\">375px</td>\n<td style=\"text-align:right\">0.5px</td>\n<td style=\"text-align:right\">375px</td>\n</tr>\n<tr>\n<td style=\"text-align:right\">768px</td>\n<td style=\"text-align:right\">1.024px</td>\n<td style=\"text-align:right\">768px</td>\n</tr>\n</tbody>\n</table>\n</div><p><code>px</code> 指 CSS 像素，不等于硬件面板上的单个物理像素。<code>rpx</code> 也不是更精细的像素单位，它只是按当前窗口宽度计算的相对长度。</p>\n<p>这个换算方式适合手机上的等比布局，但在平板和桌面小程序窗口中会持续放大。组件库如果把高度、圆角、图标和字体全部写成 <code>rpx</code>，大屏上的控件可能显得过大。</p>\n<h2>按尺寸意图选择单位</h2>\n<p>单位选择应该由尺寸的用途决定：</p>\n<div class=\"table-scroll\"><table>\n<thead>\n<tr>\n<th>场景</th>\n<th>推荐单位</th>\n<th>原因</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>页面栅格、左右留白、卡片宽度</td>\n<td><code>rpx</code></td>\n<td>需要跟随窗口宽度缩放</td>\n</tr>\n<tr>\n<td>750 宽设计稿中的页面布局</td>\n<td><code>rpx</code></td>\n<td>标注值可以直接映射到布局值</td>\n</tr>\n<tr>\n<td>细边框、分隔线</td>\n<td><code>px</code></td>\n<td>不需要随大屏持续变粗</td>\n</tr>\n<tr>\n<td>最小点击高度、固定工具栏高度</td>\n<td><code>px</code> 或受限尺寸</td>\n<td>需要保持稳定的可用性边界</td>\n</tr>\n<tr>\n<td>插画和大面积背景</td>\n<td><code>rpx</code>、百分比</td>\n<td>通常跟随容器或页面宽度变化</td>\n</tr>\n<tr>\n<td>图标和文字</td>\n<td>依据组件层级决定</td>\n<td>页面展示可缩放，基础控件更适合受控尺寸</td>\n</tr>\n</tbody>\n</table>\n</div><p>业务页面可以大量使用 <code>rpx</code>，因为页面通常直接对应设计稿。组件库需要覆盖手机、平板、分屏和不同业务密度，不能默认所有尺寸都按屏幕宽度线性增长。</p>\n<p>对大屏适配要求较高时，可以组合相对尺寸和上限：</p>\n<pre><code class=\"language-css\">.page {\n  padding-inline: min(32rpx, 24px);\n}\n\n.card {\n  width: min(686rpx, 720px);\n  margin-inline: auto;\n}\n</code></pre>\n<p>请先在目标微信基础库版本中验证 <code>min()</code> 等 CSS 函数的兼容性。如果项目需要覆盖较旧版本，可以在构建阶段生成等价样式或使用媒体查询。</p>\n<h2>translateX 可以使用 rpx</h2>\n<p>在 WXSS 中，<code>transform</code> 的长度值可以直接使用 <code>rpx</code>：</p>\n<pre><code class=\"language-css\">.panel {\n  transform: translateX(24rpx);\n  transition: transform 180ms ease;\n}\n\n.panel.is-open {\n  transform: translateX(0);\n}\n</code></pre>\n<p>搜索中常见的疑问是 <code>translateX</code> 是否支持 <code>rpx</code>。只要位移写在 WXSS 中，浏览器样式解析会处理这个长度单位。</p>\n<p>JavaScript 动画 API 接收数值时，不能假设这个数值也是 <code>rpx</code>。先按当前窗口宽度转成 <code>px</code>：</p>\n<pre><code class=\"language-js\">function rpxToPx(value) {\n  const { windowWidth } = wx.getWindowInfo();\n  return value * windowWidth / 750;\n}\n\nconst animation = wx.createAnimation({ duration: 180 });\nanimation.translateX(rpxToPx(24)).step();\n</code></pre>\n<p>窗口尺寸可能在分屏、横竖屏切换或桌面环境中变化。不要在模块加载时永久缓存换算比例，应在创建动画或窗口变化后重新计算。</p>\n<h2>自研组件库要定义尺寸契约</h2>\n<p>自研组件库可以同时使用两类设计 token：</p>\n<pre><code class=\"language-scss\">$space-page: 24rpx;\n$control-height: 44px;\n$hairline: 1px;\n$dialog-width: 640rpx;\n</code></pre>\n<p>组件文档需要说明哪些值会随屏幕缩放，哪些值保持稳定。业务方覆盖样式时，只调整公开的 token 或样式接口，不直接覆盖组件内部选择器。</p>\n<p>如果团队需要两套密度，可以在构建阶段生成 mobile 和 compact 变体。这里改变的是明确列出的 token，不是把产物中的每一个 <code>px</code> 批量替换成 <code>rpx</code>。</p>\n<h2>第三方组件库不要全量转换</h2>\n<p>第三方组件库已经基于自己的尺寸体系完成视觉和交互测试。PostCSS 对 <code>node_modules</code> 做全量换算会带来这些问题：</p>\n<ul>\n<li><code>1px</code> 边框可能被转换并出现取整差异</li>\n<li>图标、遮罩、弹层定位和动画位移一起变化</li>\n<li>组件文档中的尺寸与实际产物不再一致</li>\n<li>升级后难以区分上游改动和本地转换结果</li>\n</ul>\n<p>优先使用组件库公开的 CSS 变量、样式属性、主题配置或 wrapper 尺寸。确实需要转换时，只处理团队拥有的源码，并使用白名单明确转换目录和属性。</p>\n<p>Taro、Mpx、uni-app 等框架都有自己的样式编译规则。项目应该只保留一条单位转换链路，并把输入单位、设计稿宽度、忽略规则写进构建配置。重复转换比不转换更难排查。</p>\n<h2>按项目类型做选择</h2>\n<p>可以用下面的决策表作为默认起点：</p>\n<div class=\"table-scroll\"><table>\n<thead>\n<tr>\n<th>项目类型</th>\n<th>默认策略</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>普通业务页面</td>\n<td>布局使用 <code>rpx</code>，固定边界使用 <code>px</code></td>\n</tr>\n<tr>\n<td>自研业务组件</td>\n<td>通过 token 明确混合使用，不做隐式转换</td>\n</tr>\n<tr>\n<td>跨业务组件库</td>\n<td>以稳定尺寸为基础，为页面级间距开放 <code>rpx</code> token</td>\n</tr>\n<tr>\n<td>第三方组件库</td>\n<td>保留原单位，通过公开接口适配</td>\n</tr>\n<tr>\n<td>Taro、Mpx、uni-app 项目</td>\n<td>遵循框架编译配置，只转换自有源码</td>\n</tr>\n</tbody>\n</table>\n</div><p>最后用真实设备和窗口尺寸验证，而不是只看 375px 的开发工具预览。至少覆盖一台窄屏手机、一台主流手机和一个平板或分屏窗口，并检查文字换行、点击区域、弹层定位和动画位移。</p>\n<p>大型小程序还要把单位策略放进工程治理。滴滴出行项目使用 Mpx 处理多业务和构建问题，相关背景可以参考 <a href=\"/posts/2020/%E6%BB%B4%E6%BB%B4%E5%87%BA%E8%A1%8C%E5%B0%8F%E7%A8%8B%E5%BA%8F%E4%BD%93%E7%A7%AF%E4%BC%98%E5%8C%96%E5%AE%9E%E8%B7%B5/\">大型小程序体积治理：滴滴出行的分包、依赖与架构取舍</a>。</p>\n<h2>参考资料</h2>\n<ul>\n<li><a href=\"https://developers.weixin.qq.com/miniprogram/dev/framework/view/wxss.html\">微信开放文档：WXSS</a></li>\n<li><a href=\"https://developers.weixin.qq.com/miniprogram/dev/api/base/system/wx.getWindowInfo.html\">微信开放文档：wx.getWindowInfo</a></li>\n<li><a href=\"https://developers.weixin.qq.com/miniprogram/dev/api/ui/animation/Animation.translateX.html\">微信开放文档：Animation.translateX</a></li>\n</ul>\n","date_published":"2025-02-15T00:00:00.000Z","date_modified":"2026-08-01T00:00:00.000Z","tags":["前端","小程序"],"language":"zh"},{"id":"https://www.lihuanyu.com/en/posts/2025/deepseek-r1-sillytavern-ollama-local-deployment/","url":"https://www.lihuanyu.com/en/posts/2025/deepseek-r1-sillytavern-ollama-local-deployment/","title":"DeepSeek R1 with SillyTavern: Local Ollama Setup on Windows","summary":"A practical Windows setup for running DeepSeek R1 with Ollama and connecting SillyTavern to the local model service, covering model size, GPU requirements, API settings, and common problems.","content_html":"<p>After DeepSeek R1 became popular, two common ways of using it quickly appeared.</p>\n<p>One is to use the official website or API directly. The other is to run an open model locally, then connect it to SillyTavern for role chat.</p>\n<p>This article is about the second route: running DeepSeek R1 locally on Windows with Ollama, then connecting SillyTavern to the Ollama service on the same machine.</p>\n<p>The upside is clear: the model runs on your own computer, the setup is not too complicated, and the conversation does not have to go through a cloud model provider. The downside is just as clear: local hardware decides the experience. A small distilled model is good enough for testing and casual play, but it should not be treated as the same thing as the full online DeepSeek model.</p>\n<p>If the machine does not have a strong GPU, or if the goal is simply to use DeepSeek in SillyTavern with less friction, the API route is usually more practical: <a href=\"/en/posts/2026/deepseek-api-sillytavern-no-gpu/\">SillyTavern DeepSeek API Setup: Step-by-Step, No GPU Required</a>.</p>\n<p>For a side-by-side decision before installing either route, see <a href=\"/en/pages/sillytavern-deepseek/\">SillyTavern with DeepSeek: Local Ollama vs API Setup</a>.</p>\n<h2>What Hardware Makes Sense</h2>\n<p>The short version:</p>\n<ul>\n<li>For casual SillyTavern role chat, start with <code>deepseek-r1:8b</code> or <code>deepseek-r1:14b</code>.</li>\n<li>With about 24 GB of VRAM, <code>deepseek-r1:32b</code> becomes worth trying.</li>\n<li>Models at 70B and above are much heavier for ordinary personal machines.</li>\n<li>The official DeepSeek website and API run larger online models. A local distilled model is not the same experience.</li>\n</ul>\n<p>Ollama’s DeepSeek R1 page lists the currently available model tags and sizes. Check the official page before downloading a large file: <a href=\"https://ollama.com/library/deepseek-r1\">Ollama: deepseek-r1</a>.</p>\n<p>As of May 5, 2026, the page showed tags such as <code>8b</code>, <code>14b</code>, <code>32b</code>, <code>70b</code>, and <code>671b</code>. The <code>32b</code> model was about 20 GB, which made it a reasonable experiment for a 24 GB VRAM card.</p>\n<h2>Install SillyTavern</h2>\n<p>SillyTavern is a Web UI for role chat and character card management. It does not run a large language model by itself. Instead, it connects to providers and local backends such as OpenAI, DeepSeek, Ollama, KoboldCPP, and LM Studio.</p>\n<p>The official repository is here: <a href=\"https://github.com/SillyTavern/SillyTavern\">SillyTavern/SillyTavern</a>.</p>\n<p>On Windows, install these first:</p>\n<ul>\n<li><a href=\"https://git-scm.com/\">Git</a></li>\n<li><a href=\"https://nodejs.org/\">Node.js LTS</a></li>\n</ul>\n<p>SillyTavern’s documentation recommends the <code>release</code> branch for most users. Open a terminal in a non-system directory, such as your user directory or Documents folder, then run:</p>\n<pre><code class=\"language-bash\">git clone https://github.com/SillyTavern/SillyTavern -b release\n</code></pre>\n<p>If GitHub HTTPS access is unstable, you can configure an SSH key and clone with SSH instead:</p>\n<pre><code class=\"language-bash\">git clone git@github.com:SillyTavern/SillyTavern.git -b release\n</code></pre>\n<p>Enter the <code>SillyTavern</code> folder and double-click <code>Start.bat</code>. The first start installs dependencies. After that, a browser window usually opens automatically.</p>\n<p><img src=\"https://aipaint.lihuanyu.com/2025-02-07/%E8%BF%90%E8%A1%8C%E5%B0%8F%E9%85%92%E9%A6%86.png\" alt=\"SillyTavern after startup\"></p>\n<p>At this point, the SillyTavern interface is running, but it is not connected to any model yet.</p>\n<h2>Run DeepSeek R1 with Ollama</h2>\n<p>Ollama is a local model runner. The installation, model download, and run commands are straightforward. Download it from <a href=\"https://ollama.com/\">ollama.com</a>.</p>\n<p><img src=\"https://aipaint.lihuanyu.com/2025-02-07/ollama%E5%AE%98%E7%BD%91.png\" alt=\"Ollama website\"></p>\n<p>After installation, Windows should show the Ollama icon in the tray area, which means the local service is running. Open a terminal and type:</p>\n<pre><code class=\"language-bash\">ollama\n</code></pre>\n<p>If the command help appears, Ollama is installed correctly.</p>\n<p><img src=\"https://aipaint.lihuanyu.com/2025-02-07/ollama%E5%91%BD%E4%BB%A4%E8%A1%8C%E7%95%8C%E9%9D%A2.png\" alt=\"Ollama command line\"></p>\n<p>Next, open the DeepSeek R1 page on Ollama and choose a model size:</p>\n<p><a href=\"https://ollama.com/library/deepseek-r1\">Ollama: deepseek-r1</a></p>\n<p><img src=\"https://aipaint.lihuanyu.com/2025-02-07/%E9%80%89%E6%8B%A9R1%E7%89%88%E6%9C%AC%E5%A4%8D%E5%88%B6%E5%91%BD%E4%BB%A4.png\" alt=\"Finding DeepSeek R1 and copying the run command\"></p>\n<p>For example, to run the 32B version:</p>\n<pre><code class=\"language-bash\">ollama run deepseek-r1:32b\n</code></pre>\n<p>The first run downloads the model file. The time depends on network speed and model size. On my RTX 4090 machine with 24 GB of VRAM, <code>32b</code> was smooth enough. If your VRAM is smaller, start with <code>8b</code> or <code>14b</code>.</p>\n<p><img src=\"https://aipaint.lihuanyu.com/2025-02-07/ollama%E8%BF%90%E8%A1%8Cr1%E6%95%88%E6%9E%9C%E5%9B%BE.png\" alt=\"Running DeepSeek R1 with Ollama\"></p>\n<p>One thing is worth keeping in mind: Ollama runs open weights or distilled models locally. This is great for learning, testing, and personal experiments, but it is not the same quality tier as the official online DeepSeek models. For stable long-term use, the official API or another cloud model service may be a better default.</p>\n<h2>Connect SillyTavern to Ollama</h2>\n<p>After SillyTavern starts, the page looks roughly like this:</p>\n<p><img src=\"https://aipaint.lihuanyu.com/2025-02-07/%E5%B0%8F%E9%85%92%E9%A6%86webUI%E7%95%8C%E9%9D%A2.png\" alt=\"SillyTavern Web UI\"></p>\n<p>Click the plug icon at the top to open the API connection settings. UI labels may change between versions, but the core settings are usually:</p>\n<ul>\n<li>API type: choose the Text Completion or Chat Completion option that supports Ollama.</li>\n<li>Backend service: choose Ollama.</li>\n<li>API URL: usually <code>http://127.0.0.1:11434</code>.</li>\n<li>Model: choose or type the model you just ran, for example <code>deepseek-r1:32b</code>.</li>\n</ul>\n<p><img src=\"https://aipaint.lihuanyu.com/2025-02-07/%E9%85%8D%E7%BD%AE%E9%85%92%E9%A6%86API%E4%BD%BF%E7%94%A8%E6%9C%AC%E6%9C%BAollama%E7%9A%84deepseek.png\" alt=\"Configuring SillyTavern to use local Ollama DeepSeek\"></p>\n<p>If you see a green status indicator, a model list, or a successful test response, SillyTavern has connected to the local Ollama service. You can then choose a character card and start chatting.</p>\n<p>If no model appears, check from the terminal first:</p>\n<pre><code class=\"language-bash\">ollama list\n</code></pre>\n<p>If the target model is missing, run a small model once:</p>\n<pre><code class=\"language-bash\">ollama run deepseek-r1:8b\n</code></pre>\n<p>Confirm that it replies in the terminal, then return to SillyTavern and refresh the connection.</p>\n<h2>Common Questions</h2>\n<h3>Does DeepSeek in SillyTavern have to be local?</h3>\n<p>No. Local deployment is for people who have the hardware, want offline use, or enjoy testing local models. Without a GPU, the official API is usually simpler and more stable. The API route is covered here: <a href=\"/en/posts/2026/deepseek-api-sillytavern-no-gpu/\">DeepSeek API with SillyTavern: A No-GPU Setup</a>.</p>\n<h3>Which size should I choose: 7B, 8B, 14B, or 32B?</h3>\n<p>Choose based on VRAM and patience. Smaller models respond faster and need fewer resources, but role understanding, long context handling, and complex expression are weaker. Larger models usually feel better, but download size, VRAM usage, and waiting time all increase.</p>\n<p>For a first test, start with <code>8b</code> or <code>14b</code>. Once the whole connection works, try a larger version if the machine can handle it.</p>\n<h3>What if SillyTavern connects to Ollama but shows no model?</h3>\n<p>First confirm that the Ollama service is running and the model is actually downloaded. Run <code>ollama list</code> in the terminal. If the model is there, check that SillyTavern points to <code>http://127.0.0.1:11434</code>. Do not paste the Ollama model page URL or the GitHub repository URL into the API field.</p>\n<h3>Is local Ollama better than the official DeepSeek API?</h3>\n<p>They solve different problems.</p>\n<p>Local Ollama is more controllable, can work offline, and does not charge by token. The official API usually has better model quality, better stability, and no local hardware requirement.</p>\n<p>For casual role chat and experiments, local models are fun. For regular use where reliability matters, the API route is often easier to live with.</p>\n<h2>Other Local Tools</h2>\n<p>SillyTavern is not the only way to use Ollama. For ordinary Q&amp;A, a browser extension such as Page Assist can connect to Ollama and provide a more ChatGPT-like local interface.</p>\n<p>If you want to try more local inference tools, KoboldCPP and LM Studio are also worth looking at. Ollama is simpler. KoboldCPP and LM Studio expose more model management, UI, and parameter controls.</p>\n<h2>References</h2>\n<ul>\n<li><a href=\"https://docs.sillytavern.app/installation/windows/\">SillyTavern Windows Installation</a></li>\n<li><a href=\"https://github.com/SillyTavern/SillyTavern\">SillyTavern GitHub Repository</a></li>\n<li><a href=\"https://ollama.com/library/deepseek-r1\">Ollama: deepseek-r1</a></li>\n<li><a href=\"https://api-docs.deepseek.com/quick_start/pricing\">DeepSeek API Models &amp; Pricing</a></li>\n</ul>\n<p><a href=\"/posts/2025/%E6%9C%AC%E5%9C%B0%E9%83%A8%E7%BD%B2deepseek%E4%B8%8ESillyTavern/\">Chinese version of this article</a></p>\n","date_published":"2025-02-01T00:00:00.000Z","date_modified":"2026-07-31T00:00:00.000Z","tags":["AI","DeepSeek","SillyTavern","Ollama","Local LLM","Tutorial"],"language":"en"},{"id":"https://www.lihuanyu.com/posts/2025/%E6%9C%AC%E5%9C%B0%E9%83%A8%E7%BD%B2deepseek%E4%B8%8ESillyTavern/","url":"https://www.lihuanyu.com/posts/2025/%E6%9C%AC%E5%9C%B0%E9%83%A8%E7%BD%B2deepseek%E4%B8%8ESillyTavern/","title":"DeepSeek R1 接入 SillyTavern 小酒馆：Ollama 本地部署教程","summary":"在 Windows 上用 Ollama 运行 DeepSeek R1，并把 SillyTavern 小酒馆连接到本机模型服务，包含模型选择、显存要求、API 配置和常见问题。","content_html":"<p>DeepSeek R1 出来之后，很多人会自然想到两个玩法：一个是在网页或 API 里直接使用官方模型，另一个是把开源模型跑在本机，再接入 SillyTavern 小酒馆做角色聊天。</p>\n<p>这篇记录的是第二种方案：在 Windows 上用 Ollama 本地运行 DeepSeek R1，再让 SillyTavern 连接本机的 Ollama 服务。它的好处是不用把对话发到云端，折腾成本也不高；缺点也很明确，本地硬件决定体验，普通电脑跑小参数模型可以玩，想要接近官方满血模型的效果并不现实。</p>\n<p>如果没有较强显卡，或者只是想快速在小酒馆里用 DeepSeek，直接使用官方 API 更合适，完整方案见 <a href=\"/posts/deepseek-api-sillytavern-no-gpu/\">DeepSeek API 接入 SillyTavern：不用本地显卡的小酒馆方案</a>。</p>\n<p>还没决定走本地还是 API，可以先看 <a href=\"/pages/sillytavern-deepseek/\">DeepSeek 接入 SillyTavern 指南：本地 Ollama 与 API 怎么选</a>。</p>\n<h2>适合什么配置</h2>\n<p>先把结论放前面：</p>\n<ul>\n<li>只想体验小酒馆角色聊天，可以从 <code>deepseek-r1:8b</code> 或 <code>deepseek-r1:14b</code> 开始。</li>\n<li>显存有 24GB 左右，可以尝试 <code>deepseek-r1:32b</code>。</li>\n<li>70B 以上版本对个人电脑压力明显变大，不适合多数普通本地环境。</li>\n<li>官方 DeepSeek 网页和 API 使用的是更大规模的线上模型，本地蒸馏版不能直接等同。</li>\n</ul>\n<p>Ollama 的 DeepSeek R1 页面会列出当前可用的模型版本和体积，以官方页面为准。截至 2026 年 5 月 5 日，页面上能看到 <code>8b</code>、<code>14b</code>、<code>32b</code>、<code>70b</code>、<code>671b</code> 等版本，其中 <code>32b</code> 模型体积约 20GB，已经比较适合 24GB 显存机器尝试。</p>\n<h2>安装 SillyTavern</h2>\n<p>SillyTavern 是一个面向角色聊天和角色卡管理的 Web UI。它本身不提供大模型推理能力，而是连接到 OpenAI、DeepSeek、Ollama、KoboldCPP、LM Studio 等后端服务。</p>\n<p>官方仓库地址：<a href=\"https://github.com/SillyTavern/SillyTavern\">SillyTavern/SillyTavern</a></p>\n<p>Windows 上通常需要先安装两个基础依赖：</p>\n<ul>\n<li><a href=\"https://git-scm.com/\">Git</a></li>\n<li><a href=\"https://nodejs.org/\">Node.js LTS</a></li>\n</ul>\n<p>SillyTavern 官方文档建议普通用户使用 <code>release</code> 分支。打开命令行，找一个非系统目录的位置，例如用户目录或文档目录，然后执行：</p>\n<pre><code class=\"language-bash\">git clone https://github.com/SillyTavern/SillyTavern -b release\n</code></pre>\n<p>如果 GitHub 的 HTTPS 拉取不稳定，也可以配置 SSH key 后改用 SSH 地址：</p>\n<pre><code class=\"language-bash\">git clone git@github.com:SillyTavern/SillyTavern.git -b release\n</code></pre>\n<p>进入 <code>SillyTavern</code> 文件夹后，双击 <code>Start.bat</code>。第一次启动会安装依赖，完成后通常会自动打开浏览器。</p>\n<p><img src=\"https://aipaint.lihuanyu.com/2025-02-07/%E8%BF%90%E8%A1%8C%E5%B0%8F%E9%85%92%E9%A6%86.png\" alt=\"SillyTavern启动后的截图\"></p>\n<p>到这里，小酒馆的界面已经可以打开，但它还没有连接任何模型。接下来需要准备本机 LLM 服务。</p>\n<h2>通过 Ollama 运行 DeepSeek R1</h2>\n<p>Ollama 是一个本地模型运行工具，安装、下载模型和启动模型都比较直接。官网下载地址：<a href=\"https://ollama.com/\">ollama.com</a></p>\n<p><img src=\"https://aipaint.lihuanyu.com/2025-02-07/ollama%E5%AE%98%E7%BD%91.png\" alt=\"ollama官网\"></p>\n<p>安装完成后，Windows 右下角托盘区会出现 Ollama 图标，表示本机服务已经启动。在命令行输入：</p>\n<pre><code class=\"language-bash\">ollama\n</code></pre>\n<p>如果能看到命令帮助，说明安装正常。</p>\n<p><img src=\"https://aipaint.lihuanyu.com/2025-02-07/ollama%E5%91%BD%E4%BB%A4%E8%A1%8C%E7%95%8C%E9%9D%A2.png\" alt=\"ollama命令行界面\"></p>\n<p>接下来打开 Ollama 的 DeepSeek R1 模型页面，选择合适版本：</p>\n<p><a href=\"https://ollama.com/library/deepseek-r1\">Ollama: deepseek-r1</a></p>\n<p><img src=\"https://aipaint.lihuanyu.com/2025-02-07/%E9%80%89%E6%8B%A9R1%E7%89%88%E6%9C%AC%E5%A4%8D%E5%88%B6%E5%91%BD%E4%BB%A4.png\" alt=\"搜索找到R1选择版本复制命令\"></p>\n<p>例如使用 32B 版本：</p>\n<pre><code class=\"language-bash\">ollama run deepseek-r1:32b\n</code></pre>\n<p>第一次运行会自动下载模型文件，耗时取决于网络和模型大小。我的机器是 RTX 4090，24GB 显存，跑 <code>32b</code> 版本比较流畅；如果显存更小，建议从 <code>8b</code> 或 <code>14b</code> 开始。</p>\n<p><img src=\"https://aipaint.lihuanyu.com/2025-02-07/ollama%E8%BF%90%E8%A1%8Cr1%E6%95%88%E6%9E%9C%E5%9B%BE.png\" alt=\"ollama运行r1效果图\"></p>\n<p>需要注意，Ollama 本地运行的是开源权重或蒸馏模型。它适合学习、测试和个人玩法，但和 DeepSeek 官方网页、官方 API 上的线上模型不是同一个体验等级。真正依赖稳定效果和长时间使用的场景，优先考虑官方 API 或其它云端模型服务。</p>\n<h2>连接 SillyTavern 和 Ollama</h2>\n<p>小酒馆运行后，页面大概是这样：</p>\n<p><img src=\"https://aipaint.lihuanyu.com/2025-02-07/%E5%B0%8F%E9%85%92%E9%A6%86webUI%E7%95%8C%E9%9D%A2.png\" alt=\"小酒馆webUI运行界面\"></p>\n<p>点击顶部的插头图标进入 API 连接配置。不同版本 UI 文案可能会变化，但核心配置大致是：</p>\n<ul>\n<li>API 类型：选择 Text Completion 或 Chat Completion 中支持 Ollama 的选项。</li>\n<li>后端服务：选择 Ollama。</li>\n<li>API 地址：通常是 <code>http://127.0.0.1:11434</code>。</li>\n<li>模型：选择或填写刚才运行的模型，例如 <code>deepseek-r1:32b</code>。</li>\n</ul>\n<p><img src=\"https://aipaint.lihuanyu.com/2025-02-07/%E9%85%8D%E7%BD%AE%E9%85%92%E9%A6%86API%E4%BD%BF%E7%94%A8%E6%9C%AC%E6%9C%BAollama%E7%9A%84deepseek.png\" alt=\"配置酒馆API使用本机ollama的deepseek\"></p>\n<p>看到绿色状态、模型列表或测试连接成功，就说明 SillyTavern 已经连上本机 Ollama。之后可以选择角色卡开始聊天。</p>\n<p>如果没有看到模型，先在命令行确认：</p>\n<pre><code class=\"language-bash\">ollama list\n</code></pre>\n<p>如果列表里没有目标模型，先执行一次：</p>\n<pre><code class=\"language-bash\">ollama run deepseek-r1:8b\n</code></pre>\n<p>确认模型能在命令行正常回复，再回到 SillyTavern 刷新连接。</p>\n<h2>常见问题</h2>\n<h3>DeepSeek 小酒馆必须本地部署吗？</h3>\n<p>不必须。本地部署适合有显卡、想离线或想折腾模型的人。没有显卡时，官方 API 更省事，体验也通常更稳定。API 方案见 <a href=\"/posts/deepseek-api-sillytavern-no-gpu/\">DeepSeek API 接入 SillyTavern：不用本地显卡的小酒馆方案</a>。</p>\n<h3>7B、8B、14B、32B 应该选哪个？</h3>\n<p>按显存和耐心选。小参数模型响应更快、资源要求低，但角色理解、长上下文和复杂表达会弱一些。大参数模型效果更好，但下载体积、显存占用和等待时间都会上升。普通体验可以先从 <code>8b</code> 或 <code>14b</code> 开始，确认链路跑通后再换更大的版本。</p>\n<h3>SillyTavern 连接 Ollama 后没有模型怎么办？</h3>\n<p>先确认 Ollama 服务是否启动，再确认模型是否已经下载。命令行里执行 <code>ollama list</code>，能看到模型才说明本机存在这个模型。还要检查 SillyTavern 里的地址是否是 <code>http://127.0.0.1:11434</code>，不要把模型页面地址或 GitHub 地址填进去。</p>\n<h3>本地版本和 DeepSeek 官方 API 哪个更好？</h3>\n<p>本地版本胜在可控、隐私感更强、没有按 token 计费；官方 API 胜在效果、稳定性和硬件门槛。角色聊天如果只是娱乐和测试，本地模型很好玩；如果希望长期使用，API 方案更省心。</p>\n<h2>其它本地工具</h2>\n<p>除了 SillyTavern，普通问答也可以用 Page Assist 这类浏览器插件连接 Ollama。它更像一个本地 ChatGPT 界面，适合日常问答和简单搜索增强。</p>\n<p>如果想尝试更多本地推理工具，也可以看看 KoboldCPP 或 LM Studio。Ollama 胜在简单，KoboldCPP 和 LM Studio 在模型管理、界面和参数配置上会更丰富。</p>\n<h2>参考资料</h2>\n<ul>\n<li><a href=\"https://docs.sillytavern.app/installation/windows/\">SillyTavern Windows Installation</a></li>\n<li><a href=\"https://github.com/SillyTavern/SillyTavern\">SillyTavern GitHub Repository</a></li>\n<li><a href=\"https://ollama.com/library/deepseek-r1\">Ollama: deepseek-r1</a></li>\n<li><a href=\"https://api-docs.deepseek.com/quick_start/pricing\">DeepSeek API Models &amp; Pricing</a></li>\n</ul>\n<p><a href=\"/en/posts/2025/deepseek-r1-sillytavern-ollama-local-deployment/\">English version: DeepSeek R1 with SillyTavern: Local Ollama Setup on Windows</a></p>\n","date_published":"2025-02-01T00:00:00.000Z","date_modified":"2026-07-31T00:00:00.000Z","tags":["AI","DeepSeek","SillyTavern","小酒馆","Ollama","本地部署","教程"],"language":"zh"},{"id":"https://www.lihuanyu.com/en/posts/2025/platform-algorithms-independent-creators/","url":"https://www.lihuanyu.com/en/posts/2025/platform-algorithms-independent-creators/","title":"Platforms, Algorithms, and Creators: Why Independent Blogs Still Matter","summary":"A reflection on platform power, algorithmic distribution, creator assets, and why independent blogs still matter when most attention comes from social platforms.","content_html":"<p>Over the past few years, I wrote about several topics that looked unrelated: the rise and fall of big internet companies, algorithmic content distribution, WeChat red packet covers, and why internet startups feel harder than before.</p>\n<p>Taken together, they point to the same issue: internet entry points are increasingly concentrated, and content visibility depends more and more on platform rules and algorithmic distribution. Creators may appear to own accounts, followers, and page views, but very little of that is fully under their control.</p>\n<p>Platforms are still valuable. Without platforms, most content would never get a first audience. The problem is that platforms are good for exposure, but fragile as the only place where a creator stores content, relationships, and data.</p>\n<p><a href=\"/posts/2025/%E5%9C%A8%E5%9B%BD%E5%86%85%E7%9A%84%E5%B9%B3%E5%8F%B0%E4%BD%A0%E6%B2%A1%E6%9C%89%E7%B2%89%E4%B8%9D/\">Chinese version of this article</a></p>\n<p><img src=\"https://aipaint.lihuanyu.com/2025-02-16/civitai.png\" alt=\"platform winter\"></p>\n<h2>Followers Are Not Ownership</h2>\n<p>Most content platforms show a follower count, but a follower count is not the same as reliable reach.</p>\n<p>In the earlier internet, following an account looked more like subscribing. If a reader followed someone, the next update had a relatively stable chance of showing up. Once recommendation algorithms became the main entry point, the follow relationship still existed, but it no longer guaranteed distribution priority.</p>\n<p>This is not specific to one platform. It is the general direction of content platforms. A platform optimizes for retention, interaction, ad value, ecosystem safety, and compliance. It does not optimize for the long-term asset ownership of one creator.</p>\n<p>If an algorithm decides that unfamiliar content keeps users engaged for longer, followed content may be weakened. If the platform wants to push a new feature, traffic will move toward that feature. If moderation becomes stricter, creators have to adapt.</p>\n<p>So a more accurate description is this: having followers on a platform means the platform temporarily allows an account to reach a group of users with some probability. That probability changes, and the reasons are usually not fully transparent.</p>\n<p>This does not mean platforms are malicious. A platform has to manage massive content supply, user experience, business goals, and legal risk. It will naturally keep distribution control in its own hands. Creators need to understand that follower count is a platform metric, not a complete user relationship.</p>\n<h2>Big Companies Show How Entry Points Move</h2>\n<p>When I entered university in 2013, the default reference point for Chinese internet companies was still BAT: Baidu, Alibaba, and Tencent. A common summary at the time was that Baidu was strong in technology, Alibaba in operations, and Tencent in product.</p>\n<p>More than a decade later, mobile internet and recommendation algorithms changed many assumptions. Search no longer dominates the way it did on desktop. E-commerce competition is not only about operations. Social and content consumption have been reshaped by short video. Companies such as Douyin and Pinduoduo are often described as stronger in algorithms, traffic organization, and matching supply with demand.</p>\n<p>The point is not to predict which company will win. The more important lesson is that internet entry points move.</p>\n<p>When entry points move, everyone attached to the old entry point has to adapt. Merchants adapt to new traffic costs. Developers adapt to new platform rules. Creators adapt to new content formats.</p>\n<p>In one period, long-tail search traffic may work. In another period, titles, thumbnails, completion rate, and engagement may matter more. A platform’s power comes from its ability to redefine what becomes visible. It can encourage short video, livestreaming, image-heavy posts, or a new feature it wants to grow.</p>\n<p>If a creator binds all work to one platform entry point, the uncertainty of that entry point becomes part of the creator’s life.</p>\n<h2>Red Packet Covers: Incentives Are Not Assets</h2>\n<p>In 2023, when AI image generation became popular, I made a WeChat Official Account and a Mini Program. The algorithm gave a few posts a wave of traffic, and by the 2024 Spring Festival the account had more than 1,500 followers. WeChat gave the account 1,200 custom red packet cover quotas. At the time, it felt fresh and interesting.</p>\n<p>Before the 2025 Spring Festival, WeChat gave the account 6,000 quotas. In the end, I barely used them.</p>\n<p><img src=\"https://aipaint.lihuanyu.com/2025-01-27/5838b732c7adfe193742fb106ddfe70.png\" alt=\"custom WeChat red packet cover\"></p>\n<p>This is a useful case for understanding the relationship between platforms and creators.</p>\n<p>The red packet cover is a clever product. A user sees the cover when sending or receiving money. The creator gets brand exposure. The platform connects content, social behavior, and payment scenarios. For creators, it looks like both a traffic reward and a status signal.</p>\n<p>But it is not a creator asset.</p>\n<p>First, the quota comes from platform rules. How many quotas a creator gets, when they arrive, and what the threshold is can all change.</p>\n<p>Second, the content must pass platform review. Covers involve copyright, similarity, source material, copywriting rules, and human judgment. Even if an image is generated by AI, it can still fail review if it looks too close to existing IP or someone else’s work.</p>\n<p>Third, conversion is unstable. In 2024, I made two covers. One used generic Spring Festival imagery. The other matched the topic that previously brought traffic to the account. The second one performed better because it connected with existing reader interest. Even so, very few people ultimately followed the account through the red packet cover.</p>\n<p>By 2025, the atmosphere, review process, and cost-benefit calculation had changed. The same effort no longer felt worthwhile.</p>\n<p>That is the nature of platform incentives. They can be useful in the short term, but creators should not treat them as reusable assets. What is worth keeping is the understanding of audience interest, content process, production workflow, and lessons that can move across platforms.</p>\n<h2>Startup Winter and Creator Winter</h2>\n<p>Internet startups became harder for reasons that resemble the creator economy.</p>\n<p>The earlier internet had more low-hanging fruit. Users were growing quickly, platform rules were simpler, and a product that solved a real problem could sometimes grow naturally. Today, the internet is mature. User attention is split across many apps, acquisition costs are higher, and large companies can copy or pressure new directions more easily.</p>\n<p>Content creation follows a similar pattern. In earlier stages, content supply was lower, so consistent publishing had a better chance of being discovered. Later, the number of creators grew, feeds became crowded, and simply “keep publishing” stopped being enough.</p>\n<p>Titles, covers, pacing, topic choice, account weight, interaction rate, timing, and platform priorities all affect the result.</p>\n<p>This can make people think content quality no longer matters. That is the wrong conclusion. Quality still matters, but quality does not automatically create traffic. It is a threshold, not a guarantee.</p>\n<p>A startup cannot only believe that “a good product will naturally grow.” A creator cannot only believe that “good content will naturally find readers.” Growth has become its own problem, and platforms control many of the most important entry points.</p>\n<h2>What an Independent Blog Really Preserves</h2>\n<p>An independent blog does not magically create traffic. In many cases, it grows much more slowly than a platform account. There is no recommendation feed, no trending list, no platform campaign, and rarely a sudden viral moment.</p>\n<p>But an independent blog preserves things that platforms rarely provide.</p>\n<p><strong>Content control.</strong>\nYou decide how articles are structured, how long they remain available, whether they can be updated, whether they include links, code, or long-form context. Platforms encourage the formats that work for platform consumption. A blog can serve long-term expression.</p>\n<p><strong>Stable URLs.</strong>\nAn article link can exist for years. Citations, search engine indexing, bookmarks, and references do not depend on whether a platform still wants to distribute that post.</p>\n<p><strong>Complete context.</strong>\nPlatform content often optimizes for one post’s immediate performance. A blog is better for series, project reviews, long-term opinions, and traceable thinking.</p>\n<p><strong>Data and migration.</strong>\nMarkdown files, images, code, domains, RSS, and sitemaps can be managed by the owner. Even if the framework, server, or deployment method changes later, the content can move.</p>\n<p><strong>Search and long tail.</strong>\nPlatform content often has a short lifecycle. A blog is better for search-driven discovery. Many engineering problems, tool experiences, and personal reviews may not be algorithm-friendly, but they are useful when someone searches for them at the right time.</p>\n<p>These values are not flashy, but they are solid. Platforms provide traffic opportunities. An independent blog preserves content assets.</p>\n<h2>Use Platforms, But Change Their Role</h2>\n<p>An independent blog is not a reason to leave platforms completely. Without platforms, many posts will never be discovered for the first time.</p>\n<p>A more realistic strategy is to treat platforms as distribution channels and the blog as the content base.</p>\n<p>The workflow can be simple:</p>\n<ol>\n<li>Publish important articles on the blog as the complete version.</li>\n<li>Turn one point, one case, or one conclusion into a platform-native post.</li>\n<li>Rewrite for each platform instead of forcing the same text everywhere.</li>\n<li>Point readers back to the long-term URL whenever the platform allows it.</li>\n<li>Use RSS, email, domain names, and search so readers can find you again outside the feed.</li>\n</ol>\n<p>The point is not to be anti-platform. The point is to avoid putting all accumulated value inside a platform container. Platforms are good at expanding reach. Blogs are good at preserving judgment. A platform is a public square. A blog is a study. You meet people in the square, but you keep your work in the study.</p>\n<h2>A Reminder for Personal Writing</h2>\n<p>An independent blog is not valuable just because it exists. Its value depends on whether the content is worth preserving.</p>\n<p>For me, the most valuable posts are usually not generic tutorials. They are reviews with real context:</p>\n<ul>\n<li>Why this solution was chosen over another.</li>\n<li>What actually went wrong.</li>\n<li>What the decision depended on at the time.</li>\n<li>Which assumptions still hold up later.</li>\n<li>What I would do differently today.</li>\n</ul>\n<p>This kind of writing may not spread as easily as short-form platform content, but it remains searchable, referenceable, and reorganizable years later. It may not create high immediate traffic, but it becomes part of a public knowledge archive.</p>\n<p>That is the biggest difference between a platform account and an independent blog. A platform account shows stage-by-stage performance. A blog records a long-term trajectory.</p>\n<h2>Conclusion</h2>\n<p>On platforms, creators get exposure opportunities, not complete user relationships. They get follower numbers, not stable reach. They get campaign incentives, not portable assets.</p>\n<p>Platforms still matter. Algorithmic recommendations, social relationships, trending events, and ecosystem features can all help content reach more people. But creators should admit that these powers belong to the platform, not to the account itself.</p>\n<p>The purpose of an independent blog is to keep a controllable content base outside the platform. It does not replace platform traffic, and it does not promise fast growth. It gives long-term content a stable address, gives personal judgment continuous context, and lets readers find you again without waiting for an algorithm.</p>\n<p>Use platforms. Study algorithms. But invest in what can still remain outside them.</p>\n","date_published":"2025-02-01T00:00:00.000Z","date_modified":"2026-05-04T00:00:00.000Z","tags":["Creators","Algorithms","Platforms","Independent Blogs","Content Creation"],"language":"en"},{"id":"https://www.lihuanyu.com/posts/2025/%E5%9C%A8%E5%9B%BD%E5%86%85%E7%9A%84%E5%B9%B3%E5%8F%B0%E4%BD%A0%E6%B2%A1%E6%9C%89%E7%B2%89%E4%B8%9D/","url":"https://www.lihuanyu.com/posts/2025/%E5%9C%A8%E5%9B%BD%E5%86%85%E7%9A%84%E5%B9%B3%E5%8F%B0%E4%BD%A0%E6%B2%A1%E6%9C%89%E7%B2%89%E4%B8%9D/","title":"平台、算法与创作者：为什么还需要独立博客","summary":"从大厂兴衰、算法分发、公众号红包封面和互联网创业寒冬出发，讨论创作者为什么不能只依赖平台，以及独立博客真正能沉淀什么。","content_html":"<p>过去几年，我断断续续写过几类看起来不太相关的内容：大厂的起落、算法平台的分发逻辑、公众号红包封面的流量转化，以及互联网创业为什么越来越难。</p>\n<p>这些话题放在一起，背后其实是同一个问题：互联网的入口越来越集中，内容的可见性越来越依赖平台规则和算法分发。创作者看似拥有账号、粉丝和阅读量，但真正可控的东西并不多。</p>\n<p>平台当然有价值。没有平台，绝大多数内容根本没有冷启动的机会。问题在于，平台适合获取曝光，不适合作为唯一资产。一个创作者如果只把内容、关系和数据都放在平台里，长期看会非常被动。</p>\n<p><a href=\"/en/posts/2025/platform-algorithms-independent-creators/\">English version: Platforms, Algorithms, and Creators: Why Independent Blogs Still Matter</a></p>\n<p><img src=\"https://aipaint.lihuanyu.com/2025-02-16/civitai.png\" alt=\"平台与寒冬\"></p>\n<h2>粉丝不是关系，只是一次被授权的触达</h2>\n<p>很多内容平台都会展示粉丝数，但粉丝数并不等于稳定的触达能力。</p>\n<p>早期互联网里，关注关系更接近订阅。读者关注了一个账号，只要平台不做太多干预，更新内容就有相对稳定的机会出现在读者面前。后来推荐算法变成主入口后，关注关系仍然存在，但它不再天然等于分发优先级。</p>\n<p>这不是某一个平台的问题，而是内容平台的共同趋势。平台要优化的是用户停留、互动、广告价值和整体生态安全，而不是单个创作者的长期资产。算法认为陌生内容更能留住用户时，关注内容就会被弱化；平台要扶持某个新业务时，资源就会向新业务倾斜；审核口径收紧时，创作者也只能跟着调整。</p>\n<p>所以，在平台上拥有粉丝，更准确的说法是：平台暂时允许这个账号以某种概率触达一批用户。这个概率会变化，而且变化原因通常不完全透明。</p>\n<p>这并不意味着平台作恶。平台要对海量内容、用户体验、商业收入和合规风险负责，它必然会把“控制分发”握在自己手里。创作者真正需要意识到的是：粉丝数是平台内指标，不是完整的用户关系。</p>\n<h2>大厂起落说明入口会转移</h2>\n<p>2013 年上大学时，谈到中国互联网公司，最常听到的还是 BAT：百度、阿里、腾讯。那时有一种很流行的概括：百度强在技术，阿里强在运营，腾讯强在产品。</p>\n<p>十几年过去，移动互联网和推荐算法改变了很多判断。搜索入口不再像 PC 时代那样绝对，电商竞争不只看运营能力，社交和内容消费也被短视频重新塑形。抖音、拼多多这类公司崛起后，外界常说它们更强的是算法、流量组织和供需匹配。</p>\n<p>这不是为了判断哪家公司一定赢，而是说明一个更基础的事实：互联网入口会转移。</p>\n<p>入口转移时，依附在旧入口上的人都会被迫重新适应。商家要适应新的流量成本，开发者要适应新的平台规则，创作者也要适应新的内容形态。过去能靠搜索拿到长尾流量，后来要学会标题、封面和完播率；过去公众号推送能带来稳定阅读，后来打开率和推荐机制都变得更复杂。</p>\n<p>平台的强大之处，恰恰在于它能重新定义什么内容更容易被看见。它可以鼓励短视频，可以鼓励直播，可以鼓励图文种草，也可以把流量导向正在扶持的新功能。创作者如果把自己完全绑定在某一种平台入口上，就必须接受入口变化带来的不确定性。</p>\n<h2>红包封面的案例：激励不等于资产</h2>\n<p>2023 年，AI 绘图刚火起来时，我做过一个公众号和小程序。微信算法推了一波流量，几篇文章突然有了阅读，粉丝数在 2024 年春节前后超过 1500。微信给了 1200 个红包封面名额，这在当时还挺新鲜。</p>\n<p>2025 年春节前，微信又给了 6000 个红包封面名额，但我最后基本没有继续做。</p>\n<p><img src=\"https://aipaint.lihuanyu.com/2025-01-27/5838b732c7adfe193742fb106ddfe70.png\" alt=\"自定义微信红包封面示意图\"></p>\n<p>这个变化很适合观察平台和创作者之间的关系。</p>\n<p>红包封面本身是很聪明的产品设计。用户发红包时能展示封面，公众号或视频号可以通过封面获得曝光，平台也能让内容生态和支付场景发生连接。对创作者来说，这像是一种流量奖励，也像一种身份激励。</p>\n<p>但它并不是创作者资产。</p>\n<p>第一，名额来自平台规则。平台给多少、什么时候给、门槛怎么变，创作者只能接受。</p>\n<p>第二，内容要经过平台审核。红包封面涉及版权、相似度、素材来源、文案规范和平台判断。即使图片是 AI 新生成的，只要和已有 IP 或他人作品接近，也可能无法通过。</p>\n<p>第三，转化并不稳定。2024 年我做过两款红包封面，一款春节元素，一款和当时爆过的内容相关。后者领取和使用情况更好，因为它能连接到已有读者兴趣。但即便如此，通过红包封面最终转化成关注的用户也很少。到了 2025 年，发红包的热情、平台活动气氛和审核环境都变了，同样的事情就不再值得投入。</p>\n<p>这就是平台激励的特点：它可能短期有效，但创作者不能把它当成可长期复用的资产。真正值得沉淀的，是对用户兴趣的理解、内容方法、素材流程、案例复盘和可迁移的表达能力。</p>\n<h2>创业寒冬与创作者寒冬</h2>\n<p>互联网创业变难，和内容创作变难有相似的原因。</p>\n<p>早期互联网有很多低垂果实。用户增长快，平台规则简单，新产品只要抓住一个需求，就有机会获得自然增长。今天的互联网已经成熟，用户时间被大量应用瓜分，线上项目的获客成本更高，大公司也更容易复制或压制新方向。</p>\n<p>内容创作也是类似的。早期平台内容供给不足，一个持续更新的人更容易被看见。后来创作者越来越多，平台内容越来越拥挤，单纯“坚持输出”不再足够。标题、封面、节奏、选题、账号权重、互动率、发布时间、平台扶持方向，都会影响结果。</p>\n<p>这会让很多人误以为内容质量不重要。事实不是这样。质量仍然重要，但质量不再自动带来流量。它更像门槛，而不是保证。</p>\n<p>创业者不能只相信“做出好产品自然会增长”，创作者也不能只相信“写出好内容自然会有人看”。增长本身已经变成一个单独的问题，而平台又掌握着增长的关键入口。</p>\n<h2>独立博客真正沉淀什么</h2>\n<p>独立博客不会神奇地带来流量。甚至很多时候，它的增长比平台慢得多。没有推荐流，没有热榜，没有平台活动，也很少出现一夜爆发。</p>\n<p>但独立博客有一些平台很难提供的东西。</p>\n<p><strong>第一，内容控制权。</strong>\n文章如何组织、保留多久、是否修改、是否加链接、是否放代码、是否保留长文，都由自己决定。平台更鼓励适合平台消费的内容形态，独立博客则可以服务长期表达。</p>\n<p><strong>第二，稳定 URL。</strong>\n一篇文章的链接可以存在很多年。别人引用、搜索引擎收录、读者收藏，都不依赖某个平台是否还愿意给它分发。</p>\n<p><strong>第三，完整上下文。</strong>\n平台内容通常强调单条内容的即时表现，独立博客更适合沉淀系列文章、项目复盘、长期观点和可追溯的思考路径。</p>\n<p><strong>第四，数据和迁移能力。</strong>\nMarkdown 文件、图片、代码、域名、RSS、站点地图都可以自己管理。即使将来换框架、换服务器、换部署方式，内容资产仍然可以迁移。</p>\n<p><strong>第五，搜索和长尾。</strong>\n平台内容的生命周期常常很短，独立博客更适合被搜索长期命中。很多工程问题、工具经验和个人复盘，不一定适合算法推荐，却适合在需要时被搜索到。</p>\n<p>这些价值都不热闹，但很扎实。平台给的是流量机会，独立博客沉淀的是内容资产。</p>\n<h2>平台仍然要用，但角色要变</h2>\n<p>独立博客不是让人离开平台。完全离开平台，很多内容很难被第一次发现。</p>\n<p>更现实的策略是把平台当分发渠道，把博客当内容基地。</p>\n<p>可以这样做：</p>\n<ol>\n<li>重要文章优先写在博客，形成完整版本。</li>\n<li>平台内容只摘取其中一个观点、一个案例或一段结论。</li>\n<li>视频号、公众号、小红书、微博等平台根据各自形态改写，不强行一稿多发。</li>\n<li>平台简介、评论区或相关位置尽量引导到长期地址。</li>\n<li>通过 RSS、邮件订阅、域名和搜索，让读者能在平台之外再次找到你。</li>\n</ol>\n<p>这里的关键不是“反平台”，而是不要把全部积累放在平台容器里。平台适合扩大触达，博客适合沉淀判断。平台像广场，博客像书房。广场能遇到人，书房能留下东西。</p>\n<h2>对个人写作的提醒</h2>\n<p>独立博客也不是只要存在就有价值。它真正有价值，取决于内容本身是否值得被长期保存。</p>\n<p>对我来说，博客里最值得保留的内容往往不是通用教程，而是带有真实场景的复盘：</p>\n<ul>\n<li>为什么选择这个方案，而不是另一个方案。</li>\n<li>实际遇到了什么问题。</li>\n<li>当时的判断依据是什么。</li>\n<li>后来回看，哪些判断成立，哪些判断过时。</li>\n<li>如果今天重新做，会怎么调整。</li>\n</ul>\n<p>这类内容放在平台上，可能不如短平快内容容易传播；但放在博客里，几年后仍然能被搜索、引用和重新整理。它不一定有很高的即时流量，却能构成一个人的公开知识档案。</p>\n<p>这也是独立博客和平台账号最大的差异：平台账号展示的是阶段性表现，独立博客记录的是长期轨迹。</p>\n<h2>总结</h2>\n<p>在平台上，创作者得到的是曝光机会，不是完整的用户关系；得到的是粉丝数字，不是稳定的触达权；得到的是活动激励，不是可自由迁移的资产。</p>\n<p>平台仍然重要。算法推荐、社交关系、热点活动和平台生态，都能帮助内容被更多人看见。但创作者需要承认：这些能力属于平台，不属于账号本身。</p>\n<p>独立博客的意义，就是在平台之外保留一个可控的内容基地。它不负责替代平台流量，也不承诺快速增长。它负责让长期内容有稳定地址，让个人判断有连续上下文，让读者可以绕过算法再次找到你。</p>\n<p>所以，平台可以继续用，算法也可以继续研究。但真正值得长期经营的，是平台之外还能留下来的东西。</p>\n","date_published":"2025-02-01T00:00:00.000Z","date_modified":"2026-05-04T00:00:00.000Z","tags":["随笔","自媒体","算法","独立博客","内容创作"],"language":"zh"},{"id":"https://www.lihuanyu.com/posts/2025/%E4%BB%80%E4%B9%88%E6%98%AF%E5%89%8D%E7%AB%AF%E6%9E%B6%E6%9E%84/","url":"https://www.lihuanyu.com/posts/2025/%E4%BB%80%E4%B9%88%E6%98%AF%E5%89%8D%E7%AB%AF%E6%9E%B6%E6%9E%84/","title":"【译】什么是前端架构","summary":"翻译并整理前端架构的核心观点，强调架构不是目录结构，而是围绕业务驱动因素、权衡取舍和限制做出的重要决策。","content_html":"<blockquote>\n<p>英语原文在此：<a href=\"https://ducin.dev/what-is-frontend-architecture\">https://ducin.dev/what-is-frontend-architecture</a>\n一开始看到了其他人的翻译，比较认可这篇文章的不少内容，所以进行一个转载，但又不想纠结于一些版权方面的问题，所以干脆基于原文让最近大火的 DeepSeek R1 帮我翻译一遍。</p>\n</blockquote>\n<blockquote>\n<p>当你思考系统设计时，不要纠结于技术选型，而应聚焦于你希望系统具备的核心特性。技术选型只是这些特性的载体。 —— Gregor Hohpe</p>\n</blockquote>\n<p><strong>免责声明</strong>：如果你自认为只是个&quot;码农&quot;，请立即关闭本页面😉</p>\n<p>前端社区存在一个普遍问题😉：我们过度关注库、框架、打包工具、GitHub star 数等次要因素。我们常会狂热追捧某个工具（比如2015-2016年的Redux），然后滥用它。</p>\n<p><img src=\"https://aipaint.lihuanyu.com/2025-01-29/blog-it-javascript-frameworks.png\" alt=\"新框架新工具对前端开发者的蜜汁吸引力\"></p>\n<p>后来，我们又因同样的原因彻底厌恶这个工具（比如现在的Redux）。这种爱恨都没有道理…究竟发生了什么？🤨</p>\n<p>&quot;问题&quot;根源在于：许多前端开发者缺乏软件架构的基本认知，因为我们总把注意力放在别处。而这些架构能力恰恰是项目长期成功的关键（虽非唯一因素）。因为架构是连接业务价值和技术实现的隐形桥梁。</p>\n<p>在开展开发者培训、技术咨询或团队招聘时，我常会提问：如何理解软件架构？哪些是核心要素？如何设计稳健的系统架构？架构师的角色是什么？</p>\n<p>在继续阅读前，建议你先尝试回答这些问题😉…</p>\n<p>滴答⏰…</p>\n<p>这个问题特意设计得非常开放，以便对话者能自由表达他们认为重要的观点。我不会给出任何暗示。但当回答开头是类似&quot;（前端）架构就是如何组织目录和文件[…]&quot;时，这对我来说立即成为危险信号🟥。没错，正是最近又有人这样回答，促使我写下本文。</p>\n<p>亲爱的读者，本文旨在<strong>转变你的关注焦点</strong>：启发你从不同维度思考架构。跳脱代码仓库中的结构，摆脱具体实现方案的束缚。集中精力思考你希望系统具备哪些核心特性。从更宏观的视角，你需要哪些系统能力。摆脱工具本身的局限，转而关注它们带来的权衡取舍。最重要的是——你的业务需求如何决定必要的软件能力。</p>\n<hr>\n<h2>什么是（前端）架构？</h2>\n<p>某次推文讨论中我说😅</p>\n<p><img src=\"https://aipaint.lihuanyu.com/2025-01-29/%E7%9B%AE%E5%BD%95%E7%BB%93%E6%9E%84%E4%B8%8D%E6%98%AF%E6%9E%B6%E6%9E%84.png\" alt=\"目录结构不是架构\"></p>\n<p>在评论区被问到我对架构的定义。我的简洁回答是：</p>\n<p><strong>根据业务需求做出的、塑造当前系统且未来难以变更的决策。</strong></p>\n<p>事实上，软件架构并没有单一标准定义。我强烈推荐阅读<a href=\"https://martinfowler.com/architecture/\">Martin Fowler的软件架构指南</a>。其他替代定义包括：</p>\n<ul>\n<li>你希望在项目早期就做对的那些决策</li>\n<li>架构就是关于重要的事——无论它是什么</li>\n</ul>\n<p>注意两个关键词：<strong>重要</strong> 和 <strong>决策</strong>。</p>\n<p>在进入具体案例之前（文章后续会涉及），让我们从基础要素开始。</p>\n<hr>\n<h2>决策</h2>\n<p>在项目的整个生命周期中，我们需要做出大量决策：语言平台选择、类库选型、编程范式、代码风格（Tabs还是空格🤔）…但更重要的是：</p>\n<ul>\n<li>如何确保业务优先级得到满足？</li>\n<li>如何让数十名开发者高效协作？</li>\n<li>如何实现高频部署（包含每日、每小时甚至周五的部署）？</li>\n</ul>\n<p>如你所见，这些主题的重要性差异巨大。我们分析的角度和做出的决策具有不同的相关权重。经验越丰富的开发者，越懂得在无关紧要处节省精力——特别是当决策容易修改时。</p>\n<p>那么，如何评估某个决策是否正确？</p>\n<hr>\n<h2>驱动因素</h2>\n<p>当面临困惑时，退一步审视全局总是明智的，这包括：</p>\n<ol>\n<li>业务<strong>最高优先级</strong>是什么？</li>\n<li>需要考虑哪些<strong>限制条件</strong>？</li>\n<li>哪些目标可以<strong>妥协</strong>（可选），哪些不可退让（必需）？</li>\n</ol>\n<p><strong>架构驱动因素是迫使我们在特定项目上下文中深度探索的关键要素</strong>。它相当于项目的语境过滤器，用于判断某个理论上的优势或劣势在具体情境中是否重要。</p>\n<p>典型架构驱动因素包括：</p>\n<ul>\n<li><strong>响应时间</strong>：系统必须极速响应</li>\n<li><strong>流量承载</strong>：需要处理海量请求</li>\n<li><strong>SLA/高可用</strong>：需保持约99.99%的正常运行时间</li>\n<li><strong>组织规模</strong>：需要支持数十甚至数百名开发者协作</li>\n<li><strong>上手门槛</strong>：应便于技能较弱的开发者理解</li>\n<li><strong>上市时间</strong>：因业务需求必须快速交付功能 （备注：商业成功的关键因素之一是 上市时间-TTM-Time To Marketing。TTM是指从产生想法到向客户推出最终产品或服务的时间长度。市场发展很快，延迟的TTM可能会毁掉整个商业理念。）</li>\n</ul>\n<p>在商业环境中，这些驱动因素几乎总是存在的。你的业务代表很可能直接表达过这些需求（只是未使用&quot;驱动因素&quot;这个术语）。若不能识别这些，你将可能专注错误方向，从而大幅降低成功概率。</p>\n<p>那么，我们该如何进行系统设计，以实现这些高层次目标呢？</p>\n<hr>\n<h2>权衡取舍</h2>\n<p>必须清醒认识到：所有特性都有代价。若想让系统具备某个优点，就必须接受对应的成本。让我们扩展之前的驱动因素列表，列出可能的负面影响：</p>\n<div class=\"table-scroll\"><table>\n<thead>\n<tr>\n<th>驱动因素</th>\n<th>所需代价</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>系统必须极速响应</td>\n<td>复杂度↑，灵活性↓</td>\n</tr>\n<tr>\n<td>通过水平扩展/自动扩缩容处理海量流量</td>\n<td>需从单体架构转向分布式系统</td>\n</tr>\n<tr>\n<td>保持99.99%正常运行时间</td>\n<td>需维护金丝雀发布、蓝绿部署等高成本方案</td>\n</tr>\n<tr>\n<td>支持大规模团队协作</td>\n<td>代码重复↑，基础设施复杂度↑</td>\n</tr>\n<tr>\n<td>便于初级开发者上手</td>\n<td>无法使用团队最爱的技术栈</td>\n</tr>\n<tr>\n<td>快速交付业务功能</td>\n<td>技术债务累积↑</td>\n</tr>\n</tbody>\n</table>\n</div><p>取舍的本质在于：我们可以<strong>有意识地决定</strong>哪些可以放弃。例如：</p>\n<ul>\n<li><strong>问</strong>：99.99%可用性是否必要？<br>\n<strong>答</strong>：必要，因合同条款要求</li>\n<li><strong>问</strong>：80%测试覆盖率是否必要？<br>\n<strong>答</strong>：锦上添花，非必需，可舍弃</li>\n</ul>\n<p>制定架构决策时，<strong>必须聚焦驱动因素，同时牢记取舍代价</strong>。我们的目标是达成<strong>核心诉求</strong>，但也清楚可能需要<strong>牺牲</strong>什么。</p>\n<hr>\n<h2>限制</h2>\n<p>还有一个不言而喻的真理：并非所有事情都能实现😉。</p>\n<p>有时即使所有分析都证明某个决策正确，我们仍无法实施。外部因素可能产生冲突：</p>\n<div class=\"table-scroll\"><table>\n<thead>\n<tr>\n<th>理想决策</th>\n<th>现实阻碍</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>系统必须极速响应，但需接受复杂度↑和灵活性↓</td>\n<td>因故无法重构数据库</td>\n</tr>\n<tr>\n<td>通过水平扩展/自动扩缩容处理海量流量，需转向分布式系统</td>\n<td>因故无法更换云服务商</td>\n</tr>\n<tr>\n<td>业务需要快速交付功能，但会产生技术债务</td>\n<td>因故无法扩招开发人员</td>\n</tr>\n</tbody>\n</table>\n</div><p>这就是现实的残酷之处。为了让挑战更有趣些——我们必须接受能力受限的事实😉。但我们仍然要达成目标！</p>\n<hr>\n<h2>简要回顾</h2>\n<p>快速总结架构决策的三要素：</p>\n<ol>\n<li><strong>驱动因素</strong>：业务核心诉求</li>\n<li><strong>权衡取舍</strong>：每个决策的代价</li>\n<li><strong>现实限制</strong>：不可抗的外部约束</li>\n</ol>\n<p><strong>跳出代码层面思考</strong></p>\n<hr>\n<h2>再次提问：什么是架构？</h2>\n<p>回顾我的简易定义：<br>\n<strong>根据业务需求做出的、塑造当前系统且未来难以变更的决策。</strong></p>\n<p>现在通过具体案例区分<strong>架构决策</strong>与<strong>技术决策</strong>：</p>\n<div class=\"table-scroll\"><table>\n<thead>\n<tr>\n<th>架构决策 ✅</th>\n<th>技术决策 ❌</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>是否实施微前端（MFE）及实施方式</td>\n<td>使用webpack模块联邦或其他工具</td>\n</tr>\n<tr>\n<td>重用性与隔离性的优先级抉择</td>\n<td>是否使用barrel文件（index.js/ts）</td>\n</tr>\n<tr>\n<td>模型跨模块共享 vs ACL隔离</td>\n<td>采用类/OOP还是函数式/FP</td>\n</tr>\n<tr>\n<td>状态管理采用集中共享式 vs 分布式</td>\n<td>是否使用Redux类状态库</td>\n</tr>\n<tr>\n<td>数据获取模式：PULL vs PUSH</td>\n<td>使用Promise/async await/rxjs</td>\n</tr>\n<tr>\n<td>UI对实时数据的依赖程度</td>\n<td>选择Firebase/Supabase等BaaS</td>\n</tr>\n<tr>\n<td>客户端-服务端API契约变更权限</td>\n<td>使用GraphQL/REST/SSE协议</td>\n</tr>\n<tr>\n<td>确定核心架构驱动因素</td>\n<td>是否宣称遵循&quot;最佳实践&quot;</td>\n</tr>\n<tr>\n<td>LCP（最大内容渲染）优化是必备项还是加分项</td>\n<td>UI组件代码行数（LoC）</td>\n</tr>\n<tr>\n<td>系统单租户 vs 多租户架构</td>\n<td>认证数据存于Redux/Context/useState</td>\n</tr>\n<tr>\n<td>前端容错机制设计</td>\n<td>CI流水线强制80%测试覆盖率</td>\n</tr>\n</tbody>\n</table>\n</div><p>通过这些对比，可以清晰看到架构决策与技术决策的本质区别😅。注意上述对照展现了架构决策与技术决策的根本差异。</p>\n<hr>\n<h2>为什么目录结构不应该被视作架构？</h2>\n<p>目录结构是设计工作和引入规范的产物。它们旨在帮助我们：</p>\n<ul>\n<li>更快速地开发</li>\n<li>更安全地交付（减少破坏性变更）</li>\n</ul>\n<p>但目录结构本身不是目标，而是实现更高层次目标的手段，例如：</p>\n<ul>\n<li>通过微前端/模块化架构划分限界上下文</li>\n<li>支持独立团队间的解耦部署</li>\n<li>通过ACL（访问控制列表）隔离本地模型</li>\n</ul>\n<p>显然，目录结构可能更好地适配某个架构，也可能适配度较低。但究其本质，它只是某个概念的具象化<strong>实现</strong>——属于实现<strong>细节</strong>层面，是达成最终目标的路径。</p>\n<hr>\n<h2>目录结构无法告知我们什么</h2>\n<p>许多关键架构要素无法通过目录结构推断，包括：</p>\n<ol>\n<li><strong>是否存在‘上帝类’</strong>：无法判断模块是否真正隔离为限界上下文</li>\n<li><strong>模型复用情况</strong>：无法识别契约层与前端逻辑是否共享同一模型</li>\n<li><strong>状态管理模式</strong>：无法确认是集中式共享状态还是分布式本地状态</li>\n<li><strong>代码耦合度</strong>：无法定位问题耦合点，无法判断是否应用依赖倒置原则</li>\n<li><strong>代码内聚性</strong>：无法评估模块间的功能聚合程度</li>\n</ol>\n<p>温和地说，目录结构的重要性不足以构成架构。它可能服务于某个架构（也可能不），但本身不是架构。</p>\n<hr>\n<h2>构建正确的前端架构认知</h2>\n<p>架构不是我们😘渴望、🥰向往或😤强制执行的东西。它不会突然浮现😶🌫️，更不该直接复制前公司的成功方案🥸（即便在之前公司运行良好）。</p>\n<p><strong>架构是沟通、分析和推理的产物</strong>。如同函数根据输入产生输出，架构师的职责就是持续收集知识经验，定期运行这个&quot;输入→输出&quot;函数。</p>\n<hr>\n<h2>输入源与获取方式</h2>\n<p>架构师的核心技能是与管理层、业务方和开发团队的<strong>全方位沟通</strong>。需要收集的关键信息包括：</p>\n<h3>业务维度</h3>\n<blockquote>\n<p><strong>产品核心优势</strong>：竞争差异点所在领域</p>\n</blockquote>\n<ul>\n<li><strong>核心领域保护</strong>：核心模块/功能/团队禁止外包</li>\n<li><strong>质量红线</strong>：不可妥协的质量标准</li>\n<li><strong>领域建模</strong>：采用事件风暴等DDD实践</li>\n<li><strong>非核心领域</strong>：次要优先级</li>\n</ul>\n<h3>组织维度</h3>\n<blockquote>\n<p><strong>康威定律影响</strong>：公司结构如何决定系统交付形态</p>\n</blockquote>\n<ul>\n<li><strong>开发团队规模</strong>：部门人数与团队数量</li>\n<li><strong>产品导向程度</strong>：各团队是否100%独立负责子产品（开发→测试→部署全链路）</li>\n<li><strong>企业关联关系</strong>：关联公司、技术共享、并购等可能颠覆现有团队架构的因素</li>\n</ul>\n<h3>交付维度</h3>\n<ul>\n<li><strong>预期交付速度</strong>：方案交付时间窗与技术储备匹配度</li>\n<li><strong>部署频率要求</strong>：持续交付准备度评估</li>\n</ul>\n<hr>\n<h2>进入开发维度</h2>\n<p>在技术层面，我们需要形成以下问题链：</p>\n<h3>代码复用策略</h3>\n<ul>\n<li><strong>应鼓励/抑制多少代码复用？</strong>\n<ul>\n<li>复用越多代码量越少，但团队独立性越弱（特别是在共享模块频繁变更时，与无共享架构相比）</li>\n<li><strong>修改共享模块的后果分析</strong>（文件/组件库/制品等）：\n<ul>\n<li>是否触发重建？若需要，需重建多少模块？</li>\n<li>需要多少自动化测试？</li>\n<li>需部署多少组件？同步还是异步？</li>\n<li>总体耗时多少？</li>\n<li>共享机制引入的效率损耗？\n<ul>\n<li><strong>对系统可用性/SLA的影响评估</strong></li>\n</ul>\n</li>\n</ul>\n</li>\n</ul>\n</li>\n</ul>\n<h3>系统灵活性度量</h3>\n<ul>\n<li><strong>系统灵活性与架构测量方式</strong>\n<ul>\n<li><strong>TTM（上市时间）</strong>：预期值与当前实际值对比</li>\n<li><strong>部署频率（DF）</strong>：生产环境变更次数/单位时间</li>\n<li><strong>交付周期（CLT）</strong>：从开发者开始编码到生产部署的时间跨度</li>\n<li><strong>故障率（CFR）</strong>：变更引发故障的频率</li>\n<li><strong>恢复时间（MTTR/FDRT）</strong>：故障修复耗时</li>\n</ul>\n</li>\n</ul>\n<h3>系统稳定性保障</h3>\n<p>（超越测试覆盖率等基础指标）</p>\n<ul>\n<li><strong>故障处理机制</strong>：\n<ul>\n<li>故障发生时的标准流程？</li>\n<li>该流程的实际调用频率？😄</li>\n</ul>\n</li>\n<li><strong>同步部署分析</strong>：\n<ul>\n<li>需同步部署的模块数量与体积？</li>\n<li>是否因CI/CD配置或仓库过度拆分导致依赖项冗余构建？</li>\n</ul>\n</li>\n<li><strong>团队信任机制</strong>：\n<ul>\n<li>是否信任其他团队的交付物？</li>\n<li>采用Git-flow还是主干开发？</li>\n<li>CI/CD流程如何适配这些决策？</li>\n</ul>\n</li>\n<li><strong>容错能力验证</strong>：\n<ul>\n<li>前端是否针对后端各类故障场景进行自动化测试？</li>\n</ul>\n</li>\n<li><strong>可观测性效用评估</strong>：\n<ul>\n<li>定位前端问题的耗时？</li>\n<li>回滚操作耗时？</li>\n<li>修复问题耗时？</li>\n<li>如何/何时发现核心Web指标（LCP/FID/CLS）的回归？</li>\n</ul>\n</li>\n</ul>\n<h3>跨平台策略</h3>\n<ul>\n<li><strong>用户设备与环境</strong>：Web/原生移动端/混合方案？\n<ul>\n<li><strong>复用与分叉策略</strong>：\n<ul>\n<li>哪些组件应复用？</li>\n<li>哪些应为减少跨团队依赖而分叉？</li>\n</ul>\n</li>\n<li><strong>团队组织模式</strong>：\n<ul>\n<li>按技术平台划分 vs 按限界上下文划分？😉</li>\n</ul>\n</li>\n</ul>\n</li>\n</ul>\n<h3>特殊需求案例</h3>\n<ul>\n<li><strong>预期/必需的系统特性</strong>\n<ul>\n<li>\n<p><strong>实时协作模式</strong>：</p>\n<ul>\n<li>现状：仅支持单用户操作</li>\n<li>需求：多用户实时编辑同一数据集</li>\n<li>方案对比：\n<ul>\n<li>直接状态修改（set/update）→ 缺乏共享模型</li>\n<li>命令模式（如Redux Action）→ 天然支持协作迭代</li>\n<li>进阶方案：CRDT（无冲突复制数据类型）</li>\n</ul>\n</li>\n</ul>\n</li>\n<li>\n<p><strong>数据实时性要求</strong>：</p>\n<ul>\n<li>社交帖子点赞数延迟 → 可容忍</li>\n<li>银行系统账户余额标签切换过期 → 不可接受</li>\n<li>解决方案：\n<ul>\n<li>客户端缓存失效策略（SWR）</li>\n<li>服务端推送机制（SSE/WebSocket）</li>\n</ul>\n</li>\n</ul>\n</li>\n<li>\n<p><strong>SDK兼容性管理</strong>：</p>\n<ul>\n<li>客户基于平台SDK开发定制功能</li>\n<li>平衡法则：\n<ul>\n<li>系统演进 vs 向后兼容</li>\n<li>案例：React组件props重构\n<ul>\n<li>后果：可能引发客户重大变更</li>\n<li>困境：即使实现并测试，仍可能因&quot;无破坏性变更&quot;要求被回滚</li>\n</ul>\n</li>\n</ul>\n</li>\n</ul>\n</li>\n</ul>\n</li>\n</ul>\n<p>架构管理的本质在于提出正确问题。引用经典格言：<br>\n<strong>“我宁愿拥有无法回答的问题，也不要接受不容置疑的答案”</strong></p>\n<p>通过深入思考这些问题，我们将确定适合的架构风格与关键技术选型。</p>\n<hr>\n<h2>工具决策 vs 架构决策</h2>\n<p>有人可能会说：</p>\n<blockquote>\n<p>“嘿兄弟，但有些与工具、类库、规范或代码结构相关的决策，对项目有重大影响且后期难以更改。比如使用Redux就是架构决策！！！ 😉”</p>\n</blockquote>\n<p>嗯，是也不是。😉</p>\n<p>人们很容易执着于特定术语，却丢失了关键语境（即从业务和整体视角看什么是重要的）。</p>\n<p>确实，某些编码规范决策在多年后可能（非常）难以更改。编码规范、代码结构或目录结构（真的，任何东西）都可能加速或拖慢开发速度——当然！有些规范更合适，有些则不然。有些能让我们更快推进，有些则不能，等等。</p>\n<p>但当你跳出做出该决策的团队范围，这些就完全无关紧要了🔥。它们只是实现细节，在团队之外毫无意义。</p>\n<p><strong>示例</strong>：</p>\n<ul>\n<li>你决定每个文件只保留一个组件 → 在团队外完全无关</li>\n<li>你选择lodash/ramda工具库或不用任何库（因为&quot;非我发明&quot;） → 仍然与团队外无关</li>\n<li>你为每个模块设计特定文件结构 → 该规范影响测试、Storybook和重构 → 仍然与团队外无关<br>\n（顺便说，如果Storybook被团队外频繁使用，它就变得相关了）</li>\n</ul>\n<p>请注意：这些决策确实重要，对你的团队很关键。但仅对你的团队。它们不会带来/强制任何整体系统特性。如果决策不同，整体系统特性也不会改变。让我们进一步分析之前的说法：</p>\n<blockquote>\n<p>“使用Redux就是架构决策”</p>\n</blockquote>\n<p>（Redux对不住了😅）</p>\n<p>现在请注意：架构决策不是选择Redux本身！而是选择集中式状态管理方案，因为这可能导致模块间交叉依赖（所有人都能访问全局store的一切，对吧？），或者在将单体拆分为微前端时——使用多个独立store（如MobX）会更简单。此外，架构决策还涉及选择客户端事件溯源方案，因为业务可能需要实现实时协作功能。</p>\n<p>那么选择Redux会带来后果吗？当然。但再次强调，重点不在库本身，而在于Redux带来的高层次特性——既包括它提供的能力（前文提过），也包括引入的成本和限制。例如Redux是唯一数据源，这在考虑微前端时显然不利。Redux与其特性密不可分，但构建架构的是这些特性，而非工具本身。</p>\n<p>让我们再看一个Angular生态的例子：</p>\n<blockquote>\n<p>“不同意！如果是像NGRX这样的高层次库，选择库本身就是架构决策。需要回答多个问题：1.如何使用NGRX;2.是否总是使用Effects;3.是否通过Facade抽象;4.与哪些层级关联;5.如何跨域共享NGRX Store？”</p>\n</blockquote>\n<p>让我们一个一个来讨论：</p>\n<ul>\n<li>\n<p><strong>我们如何使用NGRX？</strong><br>\n这是个狡猾的问题，因为&quot;如何使用&quot;可能涉及高层次和低层次两个维度。模棱两可的问题😉</p>\n</li>\n<li>\n<p><strong>是否总是使用Effects？</strong><br>\n（上下文：NGRX Effects等同于redux-observable的epics——派发action后，通过rxjs响应式流处理，通常派生新action返回store）<br>\n这属于实现细节。无论选择命令式还是响应式范式，都属于编程（实现）范式，无关架构。未来可以改变这个决策。</p>\n</li>\n<li>\n<p><strong>是否通过Facade抽象？</strong><br>\n这属于封装和/或设计模式/编码模式…比架构模式低一个层级。在C4模型中属于代码层（Level 4）（实现细节）。重申——对团队重要吗？重要。对外部重要吗？不重要。</p>\n</li>\n<li>\n<p><strong>与哪些层级关联？</strong><br>\n可能涉及架构——但这与NGRX无关。使用其他状态管理方案（如React自定义hooks）时也会提出同样的问题。假设的层级（或其缺失）当然构成架构，但即使换用其他库，这个问题依然存在，对吧？</p>\n</li>\n<li>\n<p><strong>如何跨域共享NGRX Store？</strong><br>\n绝对属于架构决策。但同样与NGRX本身无关，因为使用任何其他集中式状态管理方案时都会遇到同样的问题。对吗？</p>\n</li>\n</ul>\n<p><strong>补充说明</strong>：<br>\n是否使用NGRX/redux-observables当然会影响：</p>\n<ul>\n<li>前端开发者的入门门槛</li>\n<li>他们的积极性（与工具的爱恨情仇🥹）</li>\n<li>测试编写方式等</li>\n</ul>\n<p>但重申：当你走出团队/模块/仓库范围——这些真的有那么重要吗？</p>\n<p>归根结底，决策的变更成本高低，并不决定其在大局和/或长期中的相关性。同样，在团队/仓库内部极其重要的东西，也不必然对外部具有相关性。可能有，但不必然。</p>\n<p><strong>依我拙见，是否将选择Redux称为架构决策并不重要，只要我们聚焦于该决策带来的后果。</strong></p>\n<div class=\"table-scroll\"><table>\n<thead>\n<tr>\n<th>特征</th>\n<th>工具决策</th>\n<th>架构决策</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>影响范围</td>\n<td>团队内部</td>\n<td>跨团队/系统</td>\n</tr>\n<tr>\n<td>变更成本</td>\n<td>可能高昂但局部</td>\n<td>系统性影响</td>\n</tr>\n<tr>\n<td>业务关联度</td>\n<td>间接</td>\n<td>直接驱动</td>\n</tr>\n<tr>\n<td>示例</td>\n<td>Redux/NGRX选型</td>\n<td>集中式状态管理策略</td>\n</tr>\n</tbody>\n</table>\n</div><hr>\n<h2>总结</h2>\n<p>架构的核心在于做出重要决策。这些决策应：</p>\n<ol>\n<li><strong>由业务优先级驱动</strong></li>\n<li><strong>考量权衡取舍</strong></li>\n<li><strong>适应现实限制</strong></li>\n</ol>\n<p>面对这些挑战，架构师的职责是在<strong>业务优先级/需求</strong>与<strong>技术实现/复杂度</strong>之间找到平衡点。</p>\n<p>切勿混淆以下概念：</p>\n<ul>\n<li><strong>架构</strong>：助你达成目标的高层次决策</li>\n<li><strong>实现方式</strong>：工具、类库、规范、API 等底层细节</li>\n</ul>\n<p>后者只是实现目标的可能路径，从业务优先级和现实限制的角度看，它们只是次要细节。</p>\n<p>希望本文对你有所启发，感谢阅读🤓。<br>\n特别致谢 Damian、Mateusz 和 Manfred 提供的宝贵反馈。</p>\n<blockquote>\n<p>特别特别致谢 DeepSeek R1 提供的翻译</p>\n</blockquote>\n<hr>\n","date_published":"2025-01-29T00:00:00.000Z","tags":["前端","架构","技术"],"language":"zh"},{"id":"https://www.lihuanyu.com/posts/2025/25%E5%B9%B4%E7%9A%84%E5%BE%AE%E4%BF%A1%E7%BA%A2%E5%8C%85%E5%B0%81%E9%9D%A2/","url":"https://www.lihuanyu.com/posts/2025/25%E5%B9%B4%E7%9A%84%E5%BE%AE%E4%BF%A1%E7%BA%A2%E5%8C%85%E5%B0%81%E9%9D%A2/","title":"25 年的微信红包封面","summary":"已并入《平台、算法与创作者：为什么还需要独立博客》。","content_html":"<p>关于公众号、微信红包封面、平台激励和创作者资产的复盘，已经整理进更完整的文章：</p>\n<p><a href=\"/posts/2025/%E5%9C%A8%E5%9B%BD%E5%86%85%E7%9A%84%E5%B9%B3%E5%8F%B0%E4%BD%A0%E6%B2%A1%E6%9C%89%E7%B2%89%E4%B8%9D/\">平台、算法与创作者：为什么还需要独立博客</a></p>\n<p>这页保留原链接，是因为红包封面是一个很典型的平台案例：平台给创作者提供曝光机会，也通过名额、审核、活动节奏和产品规则决定机会如何分配。</p>\n<p><img src=\"https://aipaint.lihuanyu.com/2025-01-27/5838b732c7adfe193742fb106ddfe70.png\" alt=\"自定义微信红包封面示意图\"></p>\n<p>2024 年春节，公众号粉丝数超过 1500 后，微信给过 1200 个红包封面名额。相关封面带来了一些领取、使用和访问，但最终转化成关注的用户很少。2025 年春节前，名额变成 6000 个，审核和投入产出却已经不再值得继续做。</p>\n<p><img src=\"https://aipaint.lihuanyu.com/2025-01-27/646b28b375e50d61bd89016398b6bab.png\" alt=\"24 年公众号靠 AIGC 产出的龙年红包封面\"></p>\n<p>完整文章更关注这个案例背后的结论：平台激励可以使用，但它不等于创作者资产。真正值得沉淀的，是对用户兴趣的理解、内容方法、素材流程和可迁移的表达能力。</p>\n","date_published":"2025-01-27T00:00:00.000Z","date_modified":"2026-05-04T00:00:00.000Z","tags":["微信","公众号","红包封面","AIGC"],"language":"zh"},{"id":"https://www.lihuanyu.com/posts/2024/AI%E5%BA%94%E7%94%A8%E5%BC%80%E5%8F%91%E8%80%85%E7%9A%84%E5%9B%B0%E5%B1%80/","url":"https://www.lihuanyu.com/posts/2024/AI%E5%BA%94%E7%94%A8%E5%BC%80%E5%8F%91%E8%80%85%E7%9A%84%E5%9B%B0%E5%B1%80/","title":"AI 应用开发者的困局：用户来了，账单也来了","summary":"从 AI 绘图小程序的真实数据出发，讨论独立开发者做 AI 应用时最先撞上的问题：留存低、移动端使用频率有限、推理和绘图成本真实存在，免费增长不再天然是一件好事。","content_html":"<p>AI 刚热起来的时候，很多工程师都有一种冲动：既然模型已经这么强，那是不是随手包一层产品，就能做出一个新东西？</p>\n<p>我也试过。</p>\n<p>那几年写过几个小程序，其中效果相对好的是 AI 绘图方向。不是那种惊天动地的项目，就是一个普通独立开发者能做出来的小东西：用户输入描述，后面调用模型生成图片，再围绕保存、分享、次数、广告和付费做一圈产品逻辑。</p>\n<p>从结果看，它并不算完全失败。累计用户到过 3.6 万，说明确实有人被这个东西吸引进来。可日 UV 长期只有 100 到 200，活跃用户留存大概在 10% 到 20%，新用户 7 日留存更低，经常在 1% 到 5% 之间晃。</p>\n<p>这组数字很难让人兴奋。</p>\n<p>它像一盆凉水，从头浇下来：用户是会来的，但不一定留下；功能是能跑的，但不一定成为习惯；AI 是有吸引力的，但吸引力和生意之间，还隔着很长一段路。</p>\n<h2>AI 应用很容易被试用，很难被需要</h2>\n<p>AI 应用有一种天然的演示优势。</p>\n<p>把一句话变成一张图，把一段文字改成另一种风格，把一个问题回答得头头是道。第一次看见时，很容易觉得未来已经到了。用户点进来试一把，也很自然。</p>\n<p>问题在于，试用不是需求。</p>\n<p>很多人第一次打开 AI 绘图，只是想看看它能画成什么样。画完几张，笑一下，转发一下，然后就走了。它像商场门口的抓娃娃机，围观的人不少，真正每天都来抓的人不多。</p>\n<p>工具类产品尤其容易这样。</p>\n<p>如果它没有进入用户的日常工作流，就只能靠新鲜感支撑。新鲜感是最薄的一层纸，风一吹就破。今天用户觉得 AI 绘图好玩，明天另一个平台出了视频生成，注意力就过去了。用户不是背叛了产品，只是他本来就没有把这个产品放进生活里。</p>\n<p>我自己用 AI App 也差不多。手机里装过豆包、通义千问、Kimi、元宝、Claude。后来删了一些，留下来的也不常打开。真要高频使用，更多还是在桌面端，和写作、开发、资料整理这些明确任务绑在一起。</p>\n<p>移动端不是不能做 AI，而是移动端的很多场景太碎。屏幕小，输入长文本不舒服，任务也常常不完整。用户坐在电脑前，可能真的要解决一个工作问题；用户拿着手机，多半只是等车、排队、睡前刷一下。AI 如果只是多一个聊天入口，很容易变成“想起来用一下”的东西。</p>\n<p>这对独立开发者很要命。</p>\n<p>因为独立开发者没有无限预算去教育用户，也没有足够多的产品矩阵去承接流量。一个用户进来，如果没有很快找到必须回来的理由，他就消失在人海里了。</p>\n<h2>最大的问题不是没人来，而是来了以后要花钱</h2>\n<p>传统互联网产品当然也有成本。服务器、带宽、存储、数据库，哪一样都不免费。</p>\n<p>但很多普通工具的边际成本很低。多来一个用户，多存几行数据，多打开几次页面，不至于立刻让人心疼。早期做网页、小程序、博客、后台工具，最常见的想法是：先免费放出去，看有没有人用。</p>\n<p>AI 应用不一样。</p>\n<p>用户只要真正使用模型，成本就开始发生。生成文字要 token，生成图片要 GPU，语音、视频、长上下文、Agent 任务更不用说。它不像一篇文章写好后可以被无数人阅读，更像每个用户来都要单独开一次机器。</p>\n<p>AI 绘图小程序就是这样。</p>\n<p>为了压成本，我做了弹性部署，只在有用户请求时启动 GPU，按秒计费。即便如此，一张图的成本也大概在 1 到 2 角钱。</p>\n<p>这个数字单看不大。</p>\n<p>可如果功能免费，一个用户生成 10 张图，就是 1 到 2 块钱；1000 个用户来试，就是一笔真账。用户在屏幕上点的是“生成”，开发者在背后听见的是电表声。</p>\n<p>这也是 AI 应用和普通互联网工具最不同的地方：热闹本身可能是一种危险。</p>\n<p>过去最怕没人用。现在还要怕另一件事：很多人来用，而且都是免费用。用户越活跃，账单越活跃。页面上的增长曲线往上走，云平台的扣费短信也跟着往上走。看起来像繁荣，其实可能只是亏损在加速。</p>\n<h2>免费不是不能做，但要知道谁在付钱</h2>\n<p>免费是互联网的老传统。</p>\n<p>先让用户进来，先把规模做起来，先占领心智，商业化以后再说。这套话很多时候并非错，只是它默认了一件事：规模能摊薄成本。</p>\n<p>AI 产品里，这个默认条件变弱了。</p>\n<p>如果每次有效使用都有明确成本，免费就不再只是获客策略，而是一种补贴。补贴当然可以做，但补贴要有边界。边界不清楚，产品就像路边放了一台免费咖啡机，机器越受欢迎，老板越睡不着。</p>\n<p>所以我一开始就没有考虑裸奔的 Web 形态。</p>\n<p>网页当然传播更方便，但被刷的门槛也低。AI 绘图这种功能，只要接口暴露得随便一点，很容易被人当成公共资源。小程序也有风险，但至少多了一层平台门槛，配合登录、次数、广告和付费，能把成本关在笼子里。</p>\n<p>计费也不一定意味着用户必须马上掏钱。</p>\n<p>它首先是一种限制：每天免费几次，看广告得几次，付费用户更多次数，异常用户限流。对传统小工具来说，这些东西有时显得小题大做；对 AI 应用来说，这是地基。没有这层地基，产品越好玩，越容易把自己玩死。</p>\n<p>最后这个项目勉强靠广告和少量付费打平。说赚钱，谈不上。说完全白干，也不至于。更像交了一笔学费，顺手留下一个还算能运转的小机器。</p>\n<h2>卖铲子的人往往比淘金的人舒服</h2>\n<p>AI 火起来之后，真正先赚到钱的人，未必是做应用的人。</p>\n<p>有些是 API 代理，有些是算力平台，有些是卖课，有些是包装概念、贩卖焦虑的人。应用开发者反而站在中间：上游要付模型和算力的钱，下游要说服用户付费，中间还要处理产品、工程、审核、风控、客服和增长。</p>\n<p>这很像一场淘金热。</p>\n<p>挖金子的人拿着梦想下河，卖铲子的人先把钱收了。金子可能有，也可能没有；铲子一定卖出去了。</p>\n<p>这不是说 AI 应用没有机会。恰恰相反，AI 一定会长出新的应用形态。但独立开发者最好不要被“风口”两个字冲昏头。风口只是说明天上有风，不说明地上有路。能不能走成路，还要看用户是不是真的需要，愿不愿意付费，成本能不能压住，交付质量能不能稳定。</p>\n<p>很多 AI demo 在社交媒体上很好看。放到真实产品里，就会立刻遇到一堆不浪漫的问题：</p>\n<ol>\n<li>模型偶尔失败怎么办？</li>\n<li>生成慢，用户等不等？</li>\n<li>内容不合规，责任算谁？</li>\n<li>用户反复重试，额度怎么算？</li>\n<li>高峰期排队，体验怎么保？</li>\n<li>成本涨了，价格要不要涨？</li>\n<li>便宜模型效果差，贵模型又用不起，怎么选？</li>\n</ol>\n<p>这些问题不在宣传片里，但都在账单和工单里。</p>\n<h2>模型降价会缓解问题，不会消灭问题</h2>\n<p>当然，AI 成本会下降。</p>\n<p>模型会变便宜，推理框架会优化，GPU 会换代，小模型会承担更多任务，缓存和路由也会越来越成熟。过去很贵的能力，过几年可能就变成基础设施。</p>\n<p>但成本下降不等于成本消失。</p>\n<p>带宽便宜以后，人类没有停留在文字网页，而是把图片、视频、直播、云游戏全搬了上来。存储便宜以后，人类也没有少存东西，而是拍更多照片、传更多视频、做更多备份。</p>\n<p>算力也一样。</p>\n<p>当单次生成变便宜，用户就会要求更高清、更长、更稳定、更个性化；当文本模型便宜，Agent 就会开始连续调用工具；当图片便宜，视频又会变成新胃口。技术进步会降低旧任务的成本，也会制造新任务的消耗。</p>\n<p>所以独立开发者不能只等降价。</p>\n<p>更现实的办法，是从第一天就把成本当成产品设计的一部分。简单任务用便宜模型，复杂任务再升级；能缓存就缓存，能复用就复用；失败重试要有限制；高成本功能要有明确入口；免费额度要能解释；账单要能追踪到功能和用户。</p>\n<p>这不是小气。</p>\n<p>这是 AI 应用的基本生存能力。</p>\n<h2>独立开发者要先想清楚项目性质</h2>\n<p>不是所有项目都要赚钱。</p>\n<p>练手项目、作品集、技术验证、开源实验，都可以不赚钱。它们的回报可能是经验、影响力、简历、代码资产，甚至只是开心。</p>\n<p>但如果把一个 AI 项目当产品，就不能只讲“未来”。未来太宽，账单太窄。</p>\n<p>独立开发者尤其要先想清楚几件事：</p>\n<ol>\n<li>这是一个玩具，还是一个工具？</li>\n<li>用户是偶尔试试，还是会反复使用？</li>\n<li>每次使用的成本是多少？</li>\n<li>免费用户能用到什么程度？</li>\n<li>付费理由是不是足够明确？</li>\n<li>被刷、被滥用、被薅羊毛时，系统能不能扛住？</li>\n<li>如果流量突然来了，是值得庆祝，还是要先关闸？</li>\n</ol>\n<p>这些问题想清楚，做 AI 应用就会冷静很多。</p>\n<p>AI 当然是机会。它让一个小团队甚至一个人能做出以前很难做的功能。但它同时把“生产成本”重新放回了每一次请求里。传统互联网常让人觉得软件可以无限复制，AI 则提醒开发者：每一次智能输出背后，都有真实机器在转。</p>\n<p>我后来在《<a href=\"/posts/2026/AI%E6%89%93%E7%A0%B4%E4%BA%86%E4%BA%92%E8%81%94%E7%BD%91%E7%9A%84%E9%9B%B6%E8%BE%B9%E9%99%85%E6%88%90%E6%9C%AC%E7%A5%9E%E8%AF%9D/\">AI 打破了互联网的零边际成本神话</a>》里，把这个问题又往前想了一层。AI 不只是一个新功能，它改变了软件产品的成本结构。</p>\n<p>做 AI 应用，不能只盯着模型有多聪明，也要盯着账单有多诚实。</p>\n<p>用户来了，固然是好事。</p>\n<p>但在 AI 应用里，用户来了，账单也来了。</p>\n","date_published":"2024-09-22T00:00:00.000Z","date_modified":"2026-05-16T00:00:00.000Z","tags":["AI","开发者","独立开发","商业模式"],"language":"zh"},{"id":"https://www.lihuanyu.com/posts/2024/%E5%A6%82%E4%BD%95%E6%8A%8Asvg%E6%B8%B2%E6%9F%93%E6%88%90png%E5%9B%BE%E7%89%87/","url":"https://www.lihuanyu.com/posts/2024/%E5%A6%82%E4%BD%95%E6%8A%8Asvg%E6%B8%B2%E6%9F%93%E6%88%90png%E5%9B%BE%E7%89%87/","title":"如何把svg渲染成png图片","summary":"记录把 AI 生成的 SVG 转为 PNG 的实践，比较 Cloudflare Worker、小程序 Canvas 和服务器端 sharp 方案的可行性。","content_html":"<blockquote>\n<p>简洁版：在小程序里无法把svg转为png，cloudflare 的 worker 上也不能，最终选择在自运维的服务器上转换。</p>\n</blockquote>\n<h2>背景</h2>\n<p>最近prompt大师开发了一套新的提示词很有意思，能把一个词语用鲁迅的语气，幽默、讽刺、批判性的进行解释。这个提示词要配合 Claude AI 使用，输出的内容是 SVG 。</p>\n<p>例如，对程序员这个词：\n<img src=\"https://aipaint.lihuanyu.com/2024-09-22/svg-to-png-example.png\" alt=\"\"></p>\n<h2>问题</h2>\n<p>而我们有一个微信小程序，在小程序上， svg 的展示就有一些小问题了，主要是无法预览、无法下载。</p>\n<p>那么如何能实现svg在小程序上的预览呢？ 最简单的思路当然是，转换成PNG图片。</p>\n<p>接下来的问题是，在哪转？</p>\n<h2>serverless</h2>\n<p>因为转换本身肯定是要消耗一些资源的，所以一开始是不愿意在服务器上转换的。而赛博佛祖 cloudflare 提供的 serverless 服务有非常大的免费额度，所以想着看能不能用 cloudflare 的 worker 实现这个需求。</p>\n<p>结论是不行，安装 resvg-js 后编写逻辑，运行时提示无相关能力，发现 serverless 阉割了一些底层能力，不支持 native APIs。正常写网络业务逻辑没问题，一旦需要用一些底层支持的时候就挂了。</p>\n<p>但还是不死心，搜了下相关资料 “svg to png cloudflare worker”，还真有一篇博文和一个 reddit 帖子。里面提到了一个叫 resvg-wasm 的包，通过 WebAssembly ，把 rust 编译成 wasm，以此来渲染 svg。</p>\n<p>发现确实可以，但不支持字体，在我这个场景下尤为致命。reddit 的帖子里有老哥就分享了另一个方法：svg2png-wasm，这个包支持文字。但实测发现，仅支持英文，中文不知道是我姿势不对还是字体文件太大，反正出的结果就是一堆框框。</p>\n<p>总之，折腾半天的结论就是，缺少 native APIs 的 serverless 环境不行……</p>\n<h2>小程序</h2>\n<p>那如果不能在 worker 上做，第二个思路就是，能不能在小程序里做？小程序本身是可以通过 image 标签把 svg 展示出来的，但无法预览也无法下载。</p>\n<p>那么能否结合小程序的 canvas ，把 svg 绘制在 canvas 上，再从 canvas 保存为 png 呢？</p>\n<p>答案是也不行，小程序的 canvas 不支持 svg，社区里相关问题最早在18年就出现了，但一直到24年依然是没有解决方案。不知道到底是微信的技术团队比较菜还是他们不认为这是一个高优的问题。因为 svg 的支持其实是比想象中难不少的，尤其是 svg 是可以通过引入资源等做很多复杂的事情的。</p>\n<h2>传统方案</h2>\n<p>所以花了很长的时间验证上面两条分布式的道路走不通后，最终还是妥协用传统方案来做。</p>\n<p>而传统方案的容易程度真的是震惊到我了，实在是太简单了，有系统支持下的 node ，太快乐了。</p>\n<p>安装一个 sharp 包用于解析渲染 svg ，核心代码就几行：</p>\n<pre><code class=\"language-js\">const svgBuffer = await this.downloadSvg(url);\n\n// 使用 sharp 将 SVG 转换为 PNG\nconst pngBuffer = await sharp(svgBuffer)\n  .png()\n  .toBuffer();\nres.setHeader('Content-Type', 'image/png');\nres.setHeader('Content-Disposition', 'inline; filename=&quot;converted.png&quot;');\nres.send(pngBuffer);\n</code></pre>\n<p>请求调用出图，都不用调试一次就成功了。</p>\n<p>没有压测过，但请求量一旦大了后，估计很容易崩。但没关系，这个量目测不会太大，大了再想办法解决。</p>\n<h2>字体</h2>\n<p>但仔细观察发现第一次出的图的效果，好像并不好看，至少和原作者分享的效果不一致，仔细观察了一下 svg 的内容，声明了 text 的 font-family ，标题是楷体，内容是汇文明朝体。</p>\n<p>下载字体后好看很多（就是上面放的图里的效果）。</p>\n<p>于是给服务器上也安装了对应的字体。记录一下相关过程。</p>\n<ol>\n<li>查看Ubuntu上的字体：<code>fc-list :lang=zh</code></li>\n<li>从网上下载字体或者从Windows的<code>C:\\Windows\\Fonts</code>路径把字体复制出来，注意格式是ttf，发送到服务器上。</li>\n<li>把文件复制到相应的目录<code>sudo cp -r /home/upload-font /usr/share/fonts</code></li>\n<li>执行：<code>sudo mkfontscale</code> 和 <code>sudo mkfontdir</code> 以及 <code>sudo fc-cache -fv</code> 等待字体安装</li>\n<li>再输入 <code>fc-list :lang=zh</code> 看看安装是否成功</li>\n</ol>\n<p>其实第五步里我没看到有汇文明朝体，但从svg生成的png上看确实是成功了，也没探究细节，反正实现了，先这样吧。</p>\n","date_published":"2024-09-21T00:00:00.000Z","tags":["svg"],"language":"zh"},{"id":"https://www.lihuanyu.com/posts/2024/%E5%9B%BE%E5%BA%8A%E7%B3%BB%E5%88%97%E4%B9%8Btinypng%E8%87%AA%E5%8A%A8%E5%8E%8B%E7%BC%A9%E5%9B%BE%E7%89%87/","url":"https://www.lihuanyu.com/posts/2024/%E5%9B%BE%E5%BA%8A%E7%B3%BB%E5%88%97%E4%B9%8Btinypng%E8%87%AA%E5%8A%A8%E5%8E%8B%E7%BC%A9%E5%9B%BE%E7%89%87/","title":"图床系列之 TinyPNG 自动压缩图片","summary":"已合并至 Cloudflare R2 图床完整方案。","content_html":"<p>本文已并入完整方案：</p>\n<p><a href=\"/posts/2023/%E4%BD%BF%E7%94%A8cloudflare%E6%90%AD%E5%BB%BA%E4%B8%AA%E4%BA%BA%E5%9B%BE%E5%BA%8A/\">用 Cloudflare R2 搭建个人图床：上传、压缩、访问与成本</a></p>\n<p>这部分内容只记录了在 Cloudflare Worker 里调用 TinyPNG/Tinify API 的压缩逻辑。完整方案已经把上传、R2 存储、D1 元数据、TinyPNG 压缩、查询和删除放到一个流程里。</p>\n<p>TinyPNG 压缩请求成功后，应从响应头的 <code>Location</code> 读取压缩结果地址，再请求该地址下载压缩后的图片。直接从 JSON 中读取 <code>output.url</code> 的写法并不准确。</p>\n","date_published":"2024-05-18T00:00:00.000Z","date_modified":"2026-05-03T00:00:00.000Z","tags":["图床","压缩图片"],"language":"zh"},{"id":"https://www.lihuanyu.com/posts/2024/shadcn-ui%E7%BB%84%E4%BB%B6%E5%BA%93/","url":"https://www.lihuanyu.com/posts/2024/shadcn-ui%E7%BB%84%E4%BB%B6%E5%BA%93/","title":"shadcn/ui 是什么？它为什么不是传统组件库","summary":"shadcn/ui 通过 CLI 把组件源码加入项目，而不是提供一个封装好的组件 npm 包。本文说明它的使用、升级、Registry 和团队选型方式。","content_html":"<p>shadcn/ui 是一套源码分发系统。CLI 会把组件源码和依赖加入项目，之后这份组件代码由项目自己阅读、修改和维护。它不是一个安装后只能通过公开 API 使用的传统组件 npm 包。</p>\n<p>这两种分发方式解决不同问题：</p>\n<div class=\"table-scroll\"><table>\n<thead>\n<tr>\n<th>维度</th>\n<th>shadcn/ui</th>\n<th>传统 npm 组件库</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>组件代码</td>\n<td>复制到项目中</td>\n<td>保存在依赖包中</td>\n</tr>\n<tr>\n<td>修改方式</td>\n<td>直接修改源码</td>\n<td>使用 props、主题、插槽或 wrapper</td>\n</tr>\n<tr>\n<td>升级方式</td>\n<td>查看上游差异并选择性合并</td>\n<td>更新包版本并处理 breaking changes</td>\n</tr>\n<tr>\n<td>所有权</td>\n<td>项目团队</td>\n<td>组件库维护者</td>\n</tr>\n<tr>\n<td>更适合</td>\n<td>产品差异大、需要深度改造</td>\n<td>多项目统一、集中治理</td>\n</tr>\n</tbody>\n</table>\n</div><p><a href=\"https://ui.shadcn.com/docs\">shadcn/ui 官方文档</a>把这个原则称为 Open Code：组件结构透明，并通过组合适应项目，而不是隐藏在黑盒抽象后面。</p>\n<p><img src=\"https://aipaint.lihuanyu.com/2024-02-07/shadcn-ui%E7%BB%84%E4%BB%B6%E5%BA%93demo.png\" alt=\"shadcn-ui组件库demo\"></p>\n<h2>npm 组件库解决了复用，也带来了距离</h2>\n<p>传统组件库的好处很明显。</p>\n<p>安装快，接入快，生态成熟，文档完善。一个团队不用从零写按钮、弹窗、日期选择器、表格和表单校验。尤其是后台系统，组件库就是生产工具。没有组件库，很多业务开发会退化成重复造轮子。</p>\n<p>但 npm 组件库有一种天然距离。</p>\n<p>代码在依赖包里，真正能改的是传参、插槽、样式覆盖和少量扩展点。只要需求没有越界，它很好用；一旦需求踩到组件库没准备好的地方，就开始别扭。</p>\n<p>先是覆盖样式。</p>\n<p>覆盖不了，就包一层。</p>\n<p>包一层还不够，就写一堆例外。</p>\n<p>例外多了以后，项目里长出一种很熟悉的景象：表面上使用统一组件库，实际每个复杂页面都在和组件库斗法。组件库越成熟，使用者越不敢轻易改；引用越广，维护者越不敢随便动。最后有些问题明知道是问题，也只能把它封成“兼容历史行为”。</p>\n<p>软件里有很多东西不是不能改，而是改起来牵连太大。</p>\n<p>组件库尤其如此。它本来是为了提高效率，后来也可能变成一个小型地形。所有页面都沿着它走，走得久了，路就不敢重修。</p>\n<p>这也意味着依赖管理方式发生了变化。npm 包通过版本和 lockfile 管理上游代码，shadcn/ui 则把一部分上游源码纳入当前仓库。关于前一种边界，可以参考 <a href=\"/posts/2022/%E5%89%8D%E7%AB%AF%E4%BE%9D%E8%B5%96%E4%B8%8E%E4%BF%A1%E4%BB%BB/\">前端依赖、lockfile 与可信构建</a>。</p>\n<h2>shadcn/ui 的关键是把所有权交回来</h2>\n<p>shadcn/ui 最巧的地方，不是又做了一批组件。</p>\n<p>它换了分发方式。</p>\n<p><img src=\"https://aipaint.lihuanyu.com/2024-02-07/shadcn-ui%E7%BB%84%E4%BB%B6%E5%BA%93%E4%BB%8B%E7%BB%8D.png\" alt=\"shadcn-ui组件库介绍\"></p>\n<p>用 CLI 添加一个组件，代码会进入项目目录。它通常基于 Radix 这类无样式或低样式的基础能力，再配合 Tailwind CSS、CSS variables 和一套约定好的结构。你拿到的不是黑盒组件，而是一份可以打开、阅读、修改、删减的源码。</p>\n<p>这件事很重要。</p>\n<p>因为代码一旦进了项目，责任关系就变了。</p>\n<p>传统组件库像请外面的施工队。墙怎么砌、线怎么走，大体要按别人的标准来。shadcn/ui 更像把图纸和材料交给你，房子仍然要自己盖。自由多了，责任也多了。</p>\n<p>这也是它说“不是组件库”的原因。它更像一个代码分发平台，一套组件样板，一种搭建自己组件库的起点。</p>\n<p>官方文档后来也把这个方向说得更清楚：它强调 open code、composition 和 distribution。Registry 也不只服务 React 组件，可以分发 hooks、页面、配置、规则和其他文件。</p>\n<p>这就不只是“复制粘贴组件”了。</p>\n<p>它把组件库从 npm 包，变成了一种可分发的代码资产。</p>\n<h2>CLI 如何把组件加入项目</h2>\n<p>先在现有项目中初始化配置：</p>\n<pre><code class=\"language-bash\">pnpm dlx shadcn@latest init\n</code></pre>\n<p>CLI 会创建 <code>components.json</code>，记录组件目录别名、样式、图标库和 CSS 配置。之后按需添加组件：</p>\n<pre><code class=\"language-bash\">pnpm dlx shadcn@latest add button dialog\n</code></pre>\n<p>命令会把组件文件写入项目，并安装这些组件需要的 npm 依赖。添加后应该像审查普通业务代码一样审查这些文件，而不是把它们当作生成后不可修改的产物。</p>\n<p>本地已经修改过组件时，不要直接覆盖。先预览上游变化：</p>\n<pre><code class=\"language-bash\">pnpm dlx shadcn@latest add button --dry-run\n</code></pre>\n<p>当前 CLI 还提供 <code>--diff</code> 和 <code>--view</code> 选项检查具体文件。升级的核心不是把所有组件机械同步到最新版，而是确认上游修复是否适合当前项目，再合并需要的部分。</p>\n<p>Registry 把这套分发方式扩展到了团队资产。团队可以发布自己的组件、hooks、页面区块和配置，然后仍通过 <code>shadcn add</code> 把源码加入业务仓库。它更像源码包目录，而不是运行时组件依赖。</p>\n<h2>复制不是低级，盲目复用才危险</h2>\n<p>程序员天然喜欢复用。</p>\n<p>同样的逻辑写两遍，会不舒服；同样的组件复制两份，会觉得不专业。工程训练告诉我们要抽象、要封装、要 DRY。大多数时候，这是对的。</p>\n<p>但复用不是神。</p>\n<p>复用的前提，是变化方向足够一致。如果两个地方今天长得像，明天也会一起变化，那抽成一个组件很好。如果它们只是今天看起来像，背后的业务节奏、交互细节、权限、文案、状态都不一样，强行复用就会很痛。</p>\n<p>最糟糕的复用，是把不同的东西绑在一起，然后用参数把差异一点点补回来。</p>\n<p>一开始是 <code>type</code>。</p>\n<p>后来是 <code>mode</code>。</p>\n<p>再后来是 <code>showExtra</code>、<code>enableLegacy</code>、<code>fromSpecialScene</code>。</p>\n<p>最后组件像一只塞满纸条的抽屉，什么都能放，什么都不好找。</p>\n<p>复制在这种时候反而干净。</p>\n<p>复制的好处不是偷懒，而是让变化各归各位。这个页面的按钮要特殊，就改这个页面；这个表单要多一个状态，就改这个表单；等到几个地方真的长出稳定共性，再提炼回组件。先复制，后抽象，有时比先抽象，再被抽象反噬更稳。</p>\n<p>shadcn/ui 把这个道理放到了组件分发上。</p>\n<p>它没有假设所有项目都应该永远跟着一个 npm 包走。它默认每个项目最终都会长出自己的设计系统、自己的业务口味、自己的怪需求。既然迟早要改，那不如一开始就把源码给你。</p>\n<h2>源码所有权也带来维护成本</h2>\n<p>shadcn/ui 很适合个人项目和中小项目。</p>\n<p>拿来就能用，代码看得见，改起来不心虚。做一个 SaaS、一个后台、一个内容站、一个独立产品，常见组件很快就能搭起来。它不像大而全的组件库那样带来强烈视觉气味，也不会把项目绑死在某个庞大 API 上。</p>\n<p>但它不是银弹。</p>\n<p>第一，升级成本不会消失。</p>\n<p>npm 组件库升级，至少形式上可以改版本号。虽然也可能踩坑，但路径很明确。shadcn/ui 的组件进了项目后，如果本地改过，再想跟进上游变化，就要看 diff，要判断哪些改动值得合并，哪些本地改法应该保留。CLI 能帮忙查看和应用，但不能替团队做判断。</p>\n<p>第二，团队规范要自己管。</p>\n<p>源码给了你，不代表组件库自然变好。项目里如果每个人都随手改一份 Button、Dialog、Form，最后也会变成另一种混乱。组件所有权交回来以后，团队需要更清楚地约定目录、命名、样式变量、可访问性、文档和评审规则。</p>\n<p>第三，它不替代设计能力。</p>\n<p>shadcn/ui 的默认组件好看，是因为它站在 Radix、Tailwind 和现代 Web 设计习惯上。但一个产品真正难的不是按钮圆角几像素，而是信息结构、状态设计、错误处理、权限边界、密度和可读性。组件只是木料，房子怎么住还得自己想。</p>\n<p>第四，大团队未必适合完全照搬。</p>\n<p>如果一个公司有多个产品线、统一品牌、统一设计语言、稳定设计团队和组件维护团队，以 npm 包形式发布的组件库仍然有价值。它能集中治理，统一升级，减少重复劳动，也能把设计系统当作组织资产维护。</p>\n<p>所以问题不是 AntD 错了，shadcn/ui 对了。</p>\n<p>问题是场景变了。</p>\n<h2>公司组件库仍然有意义</h2>\n<p>我以前做业务时，对组件库的感受很矛盾。</p>\n<p>没有组件库，大家各写各的，页面很快散掉；组件库太重，又容易压住业务。尤其在公司里，一个组件一旦被很多项目引用，它就不再只是代码，而是组织协作的一部分。</p>\n<p>这时 npm 包组件库有它的价值。</p>\n<p>统一版本，统一发布，统一 changelog，统一兼容策略。设计团队可以围绕它推进规范，业务团队也能用同一套语言沟通。对大型组织来说，这种中心化并不是坏事，它能减少很多无意义的分叉。</p>\n<p>但中心化最怕离业务太远。</p>\n<p>组件库维护者如果只关心抽象的优雅，不关心业务页面真实怎么用，组件就会越来越像展品。看上去端庄，拿起来硌手。业务团队为了交付，只能在外面再包一层，包着包着，公司的统一组件库就成了底座，真正好用的东西散落在各业务仓库里。</p>\n<p>shadcn/ui 提醒人的，正是这一点：</p>\n<p>组件库不只是复用问题，也是所有权问题。</p>\n<p>谁能改？谁负责？谁来判断变化该进公共层，还是留在业务层？谁承担升级成本？这些问题比“组件长什么样”更关键。</p>\n<h2>源码分发更适合 AI 阅读和改造</h2>\n<p>还有一个现在越来越明显的变化：AI 写代码时，更喜欢看得见的代码。</p>\n<p>如果组件逻辑藏在 npm 包里，AI 能看到的通常只是使用方式和类型定义。它可以帮你传参，可以帮你包一层，但很难真正理解组件细节。如果组件源码就在项目里，AI 就能读、能改、能跟着项目风格调整。</p>\n<p>这让 shadcn/ui 在 AI 时代显得更顺手。</p>\n<p>组件不再是远处的依赖，而是项目上下文的一部分。AI 可以看到 Button 怎么写，Form 怎么组织，Dialog 怎么封装，主题变量在哪里。它改出来的东西，更容易贴近项目本身。</p>\n<p>这并不意味着所有依赖都要复制进仓库。</p>\n<p>底层基础设施、复杂库、稳定协议，当然应该依赖成熟包。只是 UI 组件处在一个很特殊的位置：它离用户体验近，离业务变化近，离审美和品牌也近。越靠近这些变化，越需要可修改性。</p>\n<p>shadcn/ui 正好卡在这个位置上。</p>\n<p>它把底层复杂度交给 Radix 等成熟方案，把上层组件源码交给项目。既不完全从零造，也不完全依赖黑盒。</p>\n<h2>根据团队边界选择分发方式</h2>\n<p>如果是个人项目、小团队产品、早期 SaaS、AI 辅助开发比较多的项目，我会优先考虑 shadcn/ui 这类源码分发方案。</p>\n<p>原因很简单：改得动。</p>\n<p>早期项目最怕不是组件不够完美，而是被不适合自己的抽象绑住。源码在手里，项目可以先跑起来，再慢慢长出自己的组件层。哪怕以后不再跟上游同步，也没什么大不了，代码本来就已经是自己的了。</p>\n<p>如果是成熟公司、多产品线、大量业务共同使用一套设计体系，传统 npm 组件库仍然值得做。只是要警惕把组件库做成高高在上的东西。公共组件应该服务业务，而不是让业务围着组件库转。</p>\n<p>更现实的做法可能是混合：</p>\n<ol>\n<li>基础组件可以采用源码分发，允许项目按需改造。</li>\n<li>设计变量、图标、品牌规范和可访问性要求保持统一。</li>\n<li>业务组件不要太早上升到公共层，等多个场景真的稳定后再提炼。</li>\n<li>升级不追求机械同步，而是看变化是否解决真实问题。</li>\n<li>每个团队都要明确组件所有权，别让“复制”变成没人负责。</li>\n</ol>\n<p>shadcn/ui 的流行，不是因为它发明了按钮。</p>\n<p>它真正击中的，是很多前端项目长期存在的一种别扭：组件库提高了起步速度，也把修改权放远了。shadcn/ui 把修改权重新放回项目里，代价是团队要承担更多判断和维护责任。</p>\n<p>这笔交易很公平。</p>\n<p>代码世界里，很少有什么东西白送。传统组件库用依赖换效率，shadcn/ui 用源码换自由。选哪一个，不看口号，看项目要什么。</p>\n<p>如果一个组件未来大概率会被业务反复改造，那让它一开始就属于项目，未必是坏事。</p>\n<h2>参考资料</h2>\n<ul>\n<li><a href=\"https://ui.shadcn.com/docs\">shadcn/ui 文档</a></li>\n<li><a href=\"https://ui.shadcn.com/docs/cli\">shadcn/ui CLI</a></li>\n<li><a href=\"https://ui.shadcn.com/docs/registry\">shadcn/ui Registry</a></li>\n</ul>\n","date_published":"2024-02-07T00:00:00.000Z","date_modified":"2026-08-01T00:00:00.000Z","tags":["前端","组件库","shadcn-ui"],"language":"zh"},{"id":"https://www.lihuanyu.com/en/posts/2023/cloudflare-r2-image-hosting/","url":"https://www.lihuanyu.com/en/posts/2023/cloudflare-r2-image-hosting/","title":"Cloudflare R2 Image Hosting with Workers and D1","summary":"Build a Cloudflare R2 image host with a custom domain, authenticated Worker upload API, D1 metadata, restricted CORS, caching, and cost controls.","content_html":"<p>Use an R2 custom domain for public image delivery and keep upload, listing, and deletion behind an authenticated Worker. D1 can store searchable metadata, while a small Pages application provides the management interface.</p>\n<p>The public and management surfaces should remain separate. Readers need public access to image URLs, but only the owner should reach the Worker API or management page.</p>\n<p><a href=\"/posts/2023/%E4%BD%BF%E7%94%A8cloudflare%E6%90%AD%E5%BB%BA%E4%B8%AA%E4%BA%BA%E5%9B%BE%E5%BA%8A/\">Chinese version of this article</a></p>\n<p>This article builds a personal image hosting workflow for a static blog:</p>\n<ul>\n<li>Upload images.</li>\n<li>Compress images automatically.</li>\n<li>Generate stable public image URLs.</li>\n<li>Browse uploaded images.</li>\n<li>Copy Markdown image syntax with one action.</li>\n<li>Delete images that are no longer needed.</li>\n<li>Keep the cost low enough for personal use.</li>\n</ul>\n<h2>Choose R2 for public image delivery</h2>\n<p>Markdown keeps a static blog portable, but image storage needs its own lifecycle.</p>\n<p>Keeping images in the blog repository increases clone and deployment size over time. A small cloud server can also serve images, but it puts media traffic on the same machine as the site.</p>\n<p>Object storage is a better fit for image hosting. Cloudflare R2 is attractive here because:</p>\n<ul>\n<li>Its object storage model matches immutable image files.</li>\n<li>It supports custom domains.</li>\n<li>It does not charge egress fees, which is friendly for read-heavy personal blog traffic.</li>\n<li>It works well with Workers, D1, and Pages in the same platform.</li>\n</ul>\n<p>Cloudflare’s pricing and free quotas should be checked on the official page. This article focuses on the structure and tradeoffs that matter for a personal image host: <a href=\"https://developers.cloudflare.com/r2/pricing/\">Cloudflare R2 Pricing</a>.</p>\n<p>R2 is not the only option. Alibaba Cloud OSS, Tencent Cloud COS, and AWS S3 can all be used for similar setups. R2 is a good fit here mainly because a personal blog has mostly static image traffic, and Cloudflare’s egress policy and ecosystem match that use case well.</p>\n<h2>Separate public images from the management API</h2>\n<p>The final architecture looks like this:</p>\n<p><img src=\"https://aipaint.lihuanyu.com/2023-12-04/cloudflare%E5%9B%BE%E5%BA%8A%E6%9E%B6%E6%9E%84.jpg\" alt=\"Cloudflare personal image hosting architecture\"></p>\n<p>Cloudflare services used:</p>\n<ul>\n<li>R2: stores image files.</li>\n<li>D1: stores image metadata, such as file name, URL, created time, and size.</li>\n<li>Workers: provides upload, query, and delete APIs.</li>\n<li>Pages: hosts the frontend UI.</li>\n</ul>\n<p>Additional services:</p>\n<ul>\n<li>GitHub: stores frontend and Worker code.</li>\n<li>TinyPNG/Tinify: compresses images.</li>\n<li>Custom domain: provides long-term stable image URLs.</li>\n</ul>\n<p>Keep the trust boundary explicit:</p>\n<div class=\"table-scroll\"><table>\n<thead>\n<tr>\n<th>Surface</th>\n<th>Access</th>\n<th>Purpose</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>images.example.com</code></td>\n<td>Public</td>\n<td>Serves immutable images from the R2 custom domain</td>\n</tr>\n<tr>\n<td><code>images-admin.example.com</code></td>\n<td>Private</td>\n<td>Hosts the management UI, preferably behind Cloudflare Access</td>\n</tr>\n<tr>\n<td><code>images-api.example.com</code></td>\n<td>Private</td>\n<td>Runs the Worker upload, query, and delete API</td>\n</tr>\n</tbody>\n</table>\n</div><p>Do not embed an admin token in a static Pages bundle. For a browser-based management UI, protect the admin and API hostnames with <a href=\"https://developers.cloudflare.com/cloudflare-one/access-controls/applications/http-apps/self-hosted-public-app/\">Cloudflare Access</a>. A bearer token can remain as a second check or a local tool credential, but it must be entered at runtime and kept out of source code, build variables exposed to the browser, and <code>localStorage</code>.</p>\n<p>The finished app looks roughly like this:</p>\n<p><img src=\"https://aipaint.lihuanyu.com/2023-12-04/%E5%9B%BE%E5%BA%8A%E5%BA%94%E7%94%A8-%E5%88%97%E8%A1%A8.png\" alt=\"Image hosting app list page\"></p>\n<p><img src=\"https://aipaint.lihuanyu.com/2023-12-04/%E5%9B%BE%E5%BA%8A%E5%BA%94%E7%94%A8-%E4%B8%8A%E4%BC%A0%E5%9B%BE%E7%89%87.jpg\" alt=\"Image hosting app upload page\"></p>\n<h2>Create the R2 bucket and custom domain</h2>\n<p>Create an R2 bucket in the Cloudflare dashboard, for example:</p>\n<pre><code class=\"language-text\">image-storage\n</code></pre>\n<p>After the bucket is created, the first thing to solve is public access. R2 buckets are private by default, but an image host needs URLs that browsers can load.</p>\n<p>Cloudflare provides two options:</p>\n<ul>\n<li>Use public bucket access.</li>\n<li>Bind a custom domain.</li>\n</ul>\n<p>For a personal blog, a custom domain is the better long-term choice, for example:</p>\n<pre><code class=\"language-text\">https://aipaint.lihuanyu.com\n</code></pre>\n<p>Cloudflare documents public buckets and custom domains here: <a href=\"https://developers.cloudflare.com/r2/buckets/public-buckets/\">Public buckets and custom domains</a>.</p>\n<p>Relying on the <code>r2.dev</code> preview domain for long-term production use is risky. It is not intended as a permanent production URL, and access from mainland China may not be stable. A custom domain is a better fit for URLs that will be embedded in old posts for years.</p>\n<h2>Add D1 metadata when R2 listing is not enough</h2>\n<p>R2 can list objects by prefix and paginate with a cursor, so a basic image browser does not require another database. Add D1 when the management UI needs searchable original names, stable numeric IDs, custom ordering, or metadata that should not live on the object itself.</p>\n<p>Create a D1 database, for example:</p>\n<pre><code class=\"language-text\">image-storage-record\n</code></pre>\n<p>Table schema:</p>\n<pre><code class=\"language-sql\">CREATE TABLE IF NOT EXISTS images (\n  id INTEGER PRIMARY KEY AUTOINCREMENT,\n  object_key TEXT NOT NULL UNIQUE,\n  original_name TEXT NOT NULL,\n  image_url TEXT NOT NULL,\n  content_type TEXT,\n  size INTEGER NOT NULL DEFAULT 0,\n  created_at INTEGER NOT NULL\n);\n\nCREATE INDEX IF NOT EXISTS idx_images_created_at ON images(created_at);\n</code></pre>\n<p>Field meanings:</p>\n<ul>\n<li><code>object_key</code>: the object key in R2, for example <code>2026-05-03/uuid.png</code>.</li>\n<li><code>original_name</code>: the original uploaded file name.</li>\n<li><code>image_url</code>: the public image URL.</li>\n<li><code>content_type</code>: the image MIME type.</li>\n<li><code>size</code>: the final size stored in R2.</li>\n<li><code>created_at</code>: creation timestamp.</li>\n</ul>\n<p>For a personal image host, D1 is enough for this metadata. PostgreSQL or MySQL would add operational work without improving this workflow.</p>\n<h2>Configure Worker bindings and secrets</h2>\n<p>The Worker accesses R2 and D1 through bindings. A <code>wrangler.toml</code> can look like this:</p>\n<pre><code class=\"language-toml\">name = &quot;image-storage-worker&quot;\nmain = &quot;src/index.ts&quot;\ncompatibility_date = &quot;2026-08-01&quot;\n\n[vars]\nPUBLIC_IMAGE_BASE_URL = &quot;https://aipaint.lihuanyu.com&quot;\nADMIN_ORIGIN = &quot;https://images-admin.example.com&quot;\n\n[[r2_buckets]]\nbinding = &quot;IMAGE_BUCKET&quot;\nbucket_name = &quot;image-storage&quot;\n\n[[d1_databases]]\nbinding = &quot;DB&quot;\ndatabase_name = &quot;image-storage-record&quot;\ndatabase_id = &quot;12345678-1234-1234-1234-123456789012&quot;\n</code></pre>\n<p>The <code>database_id</code> must match the database created for the project. The <code>wrangler d1 create image-storage-record</code> command returns this value, and the dashboard also shows it on the database details page.</p>\n<p>If TinyPNG is used, the API key should be stored as a secret instead of being written into <code>wrangler.toml</code>:</p>\n<pre><code class=\"language-bash\">wrangler secret put TINIFY_API_KEY\n</code></pre>\n<p>Add an admin token as a secret when the Worker keeps the bearer-token check:</p>\n<pre><code class=\"language-bash\">wrangler secret put ADMIN_TOKEN\n</code></pre>\n<p>Worker binding configuration is documented here: <a href=\"https://developers.cloudflare.com/workers/wrangler/configuration/\">Wrangler configuration</a>.</p>\n<h2>Implement the authenticated Worker API</h2>\n<p>The following simplified Worker includes:</p>\n<ul>\n<li><code>OPTIONS</code>: handles CORS preflight requests.</li>\n<li><code>POST /upload</code>: uploads an image and optionally compresses it with TinyPNG.</li>\n<li><code>GET /query</code>: queries images with pagination.</li>\n<li><code>DELETE /delete?id=1</code>: deletes the image object and metadata.</li>\n</ul>\n<pre><code class=\"language-ts\">interface Env {\n  IMAGE_BUCKET: R2Bucket;\n  DB: D1Database;\n  PUBLIC_IMAGE_BASE_URL: string;\n  ADMIN_ORIGIN: string;\n  TINIFY_API_KEY?: string;\n  ADMIN_TOKEN?: string;\n}\n\nexport default {\n  async fetch(request: Request, env: Env): Promise&lt;Response&gt; {\n    const url = new URL(request.url);\n    const corsHeaders = createCorsHeaders(request, env);\n\n    if (!corsHeaders) {\n      return new Response('Origin not allowed', { status: 403 });\n    }\n\n    if (request.method === 'OPTIONS') {\n      return new Response(null, { headers: corsHeaders });\n    }\n\n    if (!isAuthorized(request, env)) {\n      const response = json(\n        { success: false, message: 'Unauthorized' },\n        401,\n      );\n      return withCors(response, corsHeaders);\n    }\n\n    if (request.method === 'GET' &amp;&amp; url.pathname === '/query') {\n      return withCors(await handleQuery(request, env), corsHeaders);\n    }\n\n    if (request.method === 'POST' &amp;&amp; url.pathname === '/upload') {\n      return withCors(await handleUpload(request, env), corsHeaders);\n    }\n\n    if (request.method === 'DELETE' &amp;&amp; url.pathname === '/delete') {\n      return withCors(await handleDelete(request, env), corsHeaders);\n    }\n\n    const response = json(\n      { success: false, message: 'Not found' },\n      404,\n    );\n    return withCors(response, corsHeaders);\n  },\n};\n\nfunction isAuthorized(request: Request, env: Env) {\n  const authorization = request.headers.get('Authorization');\n  return Boolean(env.ADMIN_TOKEN) &amp;&amp;\n    authorization === `Bearer ${env.ADMIN_TOKEN}`;\n}\n\nasync function handleUpload(request: Request, env: Env) {\n  const formData = await request.formData();\n  const file = formData.get('file');\n\n  if (!(file instanceof File)) {\n    return json({ success: false, message: 'Missing file' }, 400);\n  }\n\n  if (!file.type.startsWith('image/')) {\n    return json({ success: false, message: 'Only image files are allowed' }, 400);\n  }\n\n  const objectKey = createObjectKey(file.name);\n  const image = env.TINIFY_API_KEY\n    ? await compressWithTinify(file, env.TINIFY_API_KEY)\n    : {\n        body: await file.arrayBuffer(),\n        contentType: file.type || 'application/octet-stream',\n        size: file.size,\n      };\n\n  await env.IMAGE_BUCKET.put(objectKey, image.body, {\n    httpMetadata: {\n      contentType: image.contentType,\n      cacheControl: 'public, max-age=31536000, immutable',\n    },\n  });\n\n  const baseUrl = env.PUBLIC_IMAGE_BASE_URL.replace(/\\/$/, '');\n  const imageUrl = `${baseUrl}/${objectKey}`;\n  const createdAt = Date.now();\n\n  await env.DB.prepare(\n    `INSERT INTO images\n      (object_key, original_name, image_url, content_type, size, created_at)\n     VALUES (?, ?, ?, ?, ?, ?)`,\n  )\n    .bind(objectKey, file.name, imageUrl, image.contentType, image.size, createdAt)\n    .run();\n\n  return json({\n    success: true,\n    url: imageUrl,\n    markdown: `![${file.name}](${imageUrl})`,\n  });\n}\n\nasync function handleQuery(request: Request, env: Env) {\n  const url = new URL(request.url);\n  const pageNum = Math.max(Number(url.searchParams.get('pageNum')) || 1, 1);\n  const pageSize = Math.min(Math.max(Number(url.searchParams.get('pageSize')) || 20, 1), 50);\n  const offset = (pageNum - 1) * pageSize;\n\n  const list = await env.DB.prepare(\n    `SELECT id, object_key, original_name, image_url, content_type, size, created_at\n     FROM images\n     ORDER BY id DESC\n     LIMIT ? OFFSET ?`,\n  )\n    .bind(pageSize, offset)\n    .all();\n\n  const count = await env.DB.prepare(`SELECT COUNT(*) AS total FROM images`).first&lt;{\n    total: number;\n  }&gt;();\n\n  return json({\n    success: true,\n    results: list.results,\n    total: count?.total || 0,\n  });\n}\n\nasync function handleDelete(request: Request, env: Env) {\n  const url = new URL(request.url);\n  const id = Number(url.searchParams.get('id'));\n\n  if (!Number.isInteger(id) || id &lt;= 0) {\n    return json({ success: false, message: 'Invalid id' }, 400);\n  }\n\n  const row = await env.DB.prepare(`SELECT object_key FROM images WHERE id = ?`)\n    .bind(id)\n    .first&lt;{ object_key: string }&gt;();\n\n  if (!row) {\n    return json({ success: false, message: 'Image not found' }, 404);\n  }\n\n  await env.IMAGE_BUCKET.delete(row.object_key);\n  await env.DB.prepare(`DELETE FROM images WHERE id = ?`).bind(id).run();\n\n  return json({ success: true });\n}\n\nasync function compressWithTinify(file: File, apiKey: string) {\n  const source = await file.arrayBuffer();\n  const auth = `Basic ${btoa(`api:${apiKey}`)}`;\n\n  const shrink = await fetch('https://api.tinify.com/shrink', {\n    method: 'POST',\n    headers: {\n      Authorization: auth,\n      'Content-Type': file.type || 'application/octet-stream',\n    },\n    body: source,\n  });\n\n  if (!shrink.ok) {\n    const message = await shrink.text();\n    throw new Error(`TinyPNG shrink failed: ${shrink.status} ${message}`);\n  }\n\n  const outputUrl = shrink.headers.get('Location');\n\n  if (!outputUrl) {\n    throw new Error('TinyPNG did not return output location');\n  }\n\n  const optimized = await fetch(outputUrl, {\n    headers: {\n      Authorization: auth,\n    },\n  });\n\n  if (!optimized.ok) {\n    const message = await optimized.text();\n    throw new Error(`TinyPNG download failed: ${optimized.status} ${message}`);\n  }\n\n  const body = await optimized.arrayBuffer();\n\n  return {\n    body,\n    contentType: optimized.headers.get('Content-Type') || file.type || 'application/octet-stream',\n    size: Number(optimized.headers.get('Content-Length')) || body.byteLength,\n  };\n}\n\nfunction createObjectKey(filename: string) {\n  const extension = filename.includes('.') ? filename.split('.').pop() : 'bin';\n  const date = new Date().toISOString().slice(0, 10);\n  return `${date}/${crypto.randomUUID()}.${extension}`;\n}\n\nfunction createCorsHeaders(\n  request: Request,\n  env: Env,\n): Record&lt;string, string&gt; | null {\n  const origin = request.headers.get('Origin');\n\n  if (origin &amp;&amp; origin !== env.ADMIN_ORIGIN) {\n    return null;\n  }\n\n  const headers: Record&lt;string, string&gt; = {\n    'Access-Control-Allow-Methods': 'GET,POST,DELETE,OPTIONS',\n    'Access-Control-Allow-Headers': 'Content-Type,Authorization',\n    Vary: 'Origin',\n  };\n\n  if (origin) {\n    headers['Access-Control-Allow-Origin'] = origin;\n  }\n\n  return headers;\n}\n\nfunction withCors(response: Response, headers: Record&lt;string, string&gt;) {\n  const wrapped = new Response(response.body, response);\n\n  for (const [name, value] of Object.entries(headers)) {\n    wrapped.headers.set(name, value);\n  }\n\n  return wrapped;\n}\n\nfunction json(data: unknown, status = 200) {\n  return new Response(JSON.stringify(data), {\n    status,\n    headers: {\n      'Content-Type': 'application/json; charset=utf-8',\n    },\n  });\n}\n</code></pre>\n<p>There are a few details worth calling out.</p>\n<p>First, the upload API should verify <code>image/*</code>. Otherwise the image host can accidentally become arbitrary file storage.</p>\n<p>Second, the API now denies management requests when <code>ADMIN_TOKEN</code> is missing. A local client or a management UI that asks for the token at runtime can send:</p>\n<pre><code class=\"language-text\">Authorization: Bearer your_admin_token_here\n</code></pre>\n<p>Do not compile that value into frontend JavaScript. Cloudflare Access is the better browser-facing boundary because it authenticates the operator before the request reaches the Worker.</p>\n<p>Third, TinyPNG’s API does not return <code>output.url</code> in the JSON response from the shrink request. After the compression request succeeds, read the <code>Location</code> response header, then request that URL to download the optimized image. The API behavior is documented here: <a href=\"https://tinypng.com/developers/reference\">Tinify API reference</a>.</p>\n<p>Fourth, using the original file name as <code>object_key</code> causes problems with non-ASCII names, spaces, and overwrites. A date prefix plus UUID is more robust.</p>\n<p>Fifth, R2 and D1 do not share a transaction. If the D1 insert fails after an R2 upload, delete the new object or record it for reconciliation. Apply the same rule when a delete succeeds in one service but fails in the other.</p>\n<h2>Build a focused management interface</h2>\n<p>Any frontend framework works. The example implementation used SolidJS, but React, Vue, or Svelte can implement the same small set of management actions.</p>\n<p>The core functions are upload, query, and delete.</p>\n<p>Upload:</p>\n<pre><code class=\"language-ts\">async function uploadImage(file: File) {\n  const formData = new FormData();\n  formData.append('file', file);\n\n  const response = await fetch(`${apiBaseUrl}/upload`, {\n    method: 'POST',\n    headers: {\n      Authorization: `Bearer ${adminToken}`,\n    },\n    body: formData,\n  });\n\n  if (!response.ok) {\n    throw new Error(await response.text());\n  }\n\n  return response.json();\n}\n</code></pre>\n<p>Query:</p>\n<pre><code class=\"language-ts\">async function queryImages(pageNum = 1, pageSize = 20) {\n  const response = await fetch(\n    `${apiBaseUrl}/query?pageNum=${pageNum}&amp;pageSize=${pageSize}`,\n    {\n      headers: {\n        Authorization: `Bearer ${adminToken}`,\n      },\n    },\n  );\n\n  if (!response.ok) {\n    throw new Error(await response.text());\n  }\n\n  return response.json();\n}\n</code></pre>\n<p>Delete:</p>\n<pre><code class=\"language-ts\">async function deleteImage(id: number) {\n  const response = await fetch(`${apiBaseUrl}/delete?id=${id}`, {\n    method: 'DELETE',\n    headers: {\n      Authorization: `Bearer ${adminToken}`,\n    },\n  });\n\n  if (!response.ok) {\n    throw new Error(await response.text());\n  }\n\n  return response.json();\n}\n</code></pre>\n<p>The UI needs only a few interactions:</p>\n<ul>\n<li>Select or drag an image file.</li>\n<li>Show the image URL and Markdown after upload.</li>\n<li>List thumbnails, original file names, created times, and sizes.</li>\n<li>Copy URL.</li>\n<li>Copy Markdown.</li>\n<li>Delete an image.</li>\n</ul>\n<p>The frontend can be deployed to Cloudflare Pages. It can live in the same repository as the Worker API or in a separate repository. For a personal project, keeping them separate is often clearer: frontend issues do not affect image access, and the Worker API can be maintained independently.</p>\n<h2>Keep image URLs stable and cacheable</h2>\n<p>For an image host, URL stability matters most. Once an image URL is written into a post, it should not change casually.</p>\n<p>A practical setup:</p>\n<ul>\n<li>Use a separate subdomain for images, such as <code>aipaint.lihuanyu.com</code>.</li>\n<li>Bind the R2 bucket to that subdomain.</li>\n<li>Use only that subdomain in blog posts.</li>\n<li>Avoid embedding Worker preview domains or Pages preview domains in posts.</li>\n</ul>\n<p>The Worker example stores UUID-based object keys with <code>Cache-Control: public, max-age=31536000, immutable</code>. Do not replace content at one of these URLs. Upload a new object and update the post when an image changes.</p>\n<h2>Estimate R2 cost before adding other services</h2>\n<p>Cloudflare’s published R2 Standard rates on August 1, 2026 are:</p>\n<div class=\"table-scroll\"><table>\n<thead>\n<tr>\n<th>Item</th>\n<th style=\"text-align:right\">Included each month</th>\n<th style=\"text-align:right\">Standard rate after the free tier</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Storage</td>\n<td style=\"text-align:right\">10 GB-month</td>\n<td style=\"text-align:right\">$0.015 per GB-month</td>\n</tr>\n<tr>\n<td>Class A operations</td>\n<td style=\"text-align:right\">1 million</td>\n<td style=\"text-align:right\">$4.50 per million requests</td>\n</tr>\n<tr>\n<td>Class B operations</td>\n<td style=\"text-align:right\">10 million</td>\n<td style=\"text-align:right\">$0.36 per million requests</td>\n</tr>\n<tr>\n<td>Internet egress</td>\n<td style=\"text-align:right\">Free</td>\n<td style=\"text-align:right\">Free</td>\n</tr>\n</tbody>\n</table>\n</div><p>Cloudflare rounds billable usage up to the next billing unit. The free tier applies to Standard storage, not Infrequent Access storage. Check the <a href=\"https://developers.cloudflare.com/r2/pricing/\">current R2 pricing</a> before relying on these figures.</p>\n<p>The complete workflow can also incur costs from:</p>\n<ul>\n<li>R2 storage and requests.</li>\n<li>D1 reads and writes.</li>\n<li>Workers requests.</li>\n<li>TinyPNG compression usage.</li>\n</ul>\n<p>For a personal blog below the free-tier limits, R2 storage and operations can remain free. TinyPNG needs a separate calculation because it is not a Cloudflare service, and its quota follows Tinify’s own rules. Compression can also run locally before upload.</p>\n<p>The practical tradeoff is:</p>\n<ul>\n<li>R2 is a good place to store images long term.</li>\n<li>D1 only stores metadata, so its cost is negligible for this use case.</li>\n<li>Workers is well suited for lightweight APIs like this.</li>\n<li>TinyPNG is useful, but not required for the first version.</li>\n</ul>\n<p>For a writing workflow, the first version can skip TinyPNG and focus on upload, query, and copying Markdown. Compression can be added later when image volume or page load time starts to matter.</p>\n<h2>Add controls before supporting other users</h2>\n<p>This setup is suitable for a personal image host. It should not be exposed as a public platform without more work.</p>\n<p>If it is opened to other users, at least these parts are needed:</p>\n<ul>\n<li>User accounts.</li>\n<li>Permission isolation.</li>\n<li>Upload rate limits.</li>\n<li>File size limits.</li>\n<li>Content safety checks.</li>\n<li>Storage quotas.</li>\n<li>Delete audit logs.</li>\n<li>Hotlink protection or access control.</li>\n</ul>\n<p>For personal use, the most important point is to keep the upload API protected. Otherwise it can be abused as public file storage.</p>\n<h2>Build the smallest useful version first</h2>\n<p>Cloudflare R2 is a good fit for a personal blog image host, but stopping at “dashboard upload plus manually assembled URL” leaves too much friction in the writing workflow. A useful image host should connect upload, compression, list view, copy, and delete.</p>\n<p>A reasonable implementation order is:</p>\n<ol>\n<li>Create an R2 bucket and bind a custom image domain.</li>\n<li>Add a Worker upload API.</li>\n<li>Store image metadata in D1.</li>\n<li>Build a small frontend page.</li>\n<li>Add TinyPNG or another compression step later.</li>\n</ol>\n<p>This keeps the stability of object storage while making image insertion smooth enough for regular blogging.</p>\n<h2>References</h2>\n<ul>\n<li><a href=\"https://developers.cloudflare.com/r2/pricing/\">Cloudflare R2 Pricing</a></li>\n<li><a href=\"https://developers.cloudflare.com/r2/buckets/public-buckets/\">Cloudflare R2 Public buckets</a></li>\n<li><a href=\"https://developers.cloudflare.com/workers/wrangler/configuration/\">Cloudflare Workers Wrangler configuration</a></li>\n<li><a href=\"https://developers.cloudflare.com/cloudflare-one/access-controls/applications/http-apps/self-hosted-public-app/\">Cloudflare Access for self-hosted applications</a></li>\n<li><a href=\"https://tinypng.com/developers/reference\">Tinify API reference</a></li>\n</ul>\n","date_published":"2023-12-04T00:00:00.000Z","date_modified":"2026-08-01T00:00:00.000Z","tags":["Image Hosting","Cloudflare","R2","D1","Worker","TinyPNG"],"language":"en"},{"id":"https://www.lihuanyu.com/posts/2023/%E5%A4%A7%E5%8E%82%E7%9A%84%E8%B5%B7%E8%B5%B7%E8%90%BD%E8%90%BD/","url":"https://www.lihuanyu.com/posts/2023/%E5%A4%A7%E5%8E%82%E7%9A%84%E8%B5%B7%E8%B5%B7%E8%90%BD%E8%90%BD/","title":"大厂的起起落落","summary":"已并入《平台、算法与创作者：为什么还需要独立博客》。","content_html":"<p>关于 BAT、拼多多、抖音、平台入口和算法分发的判断，已经整理进更完整的文章：</p>\n<p><a href=\"/posts/2025/%E5%9C%A8%E5%9B%BD%E5%86%85%E7%9A%84%E5%B9%B3%E5%8F%B0%E4%BD%A0%E6%B2%A1%E6%9C%89%E7%B2%89%E4%B8%9D/\">平台、算法与创作者：为什么还需要独立博客</a></p>\n<p>这页保留原链接，是因为“大厂起落”仍然是理解平台权力变化的一个入口。搜索、运营、产品、算法都曾经在不同阶段代表互联网入口的主要组织方式。入口变化时，依附在入口上的商家、开发者和创作者都要重新适应。</p>\n<p><img src=\"https://aipaint.lihuanyu.com/2023-12-04/AI-%E5%B7%A1%E6%B4%8B%E8%88%B03.png\" alt=\"AI绘图\"></p>\n<p>完整文章更关注这个问题对个人创作者的影响：当内容可见性越来越依赖平台规则和算法分发时，为什么仍然需要一个可长期沉淀内容的独立博客。</p>\n","date_published":"2023-12-04T00:00:00.000Z","date_modified":"2026-05-04T00:00:00.000Z","tags":["随笔","大厂"],"language":"zh"},{"id":"https://www.lihuanyu.com/posts/2023/%E4%BD%BF%E7%94%A8cloudflare%E6%90%AD%E5%BB%BA%E4%B8%AA%E4%BA%BA%E5%9B%BE%E5%BA%8A/","url":"https://www.lihuanyu.com/posts/2023/%E4%BD%BF%E7%94%A8cloudflare%E6%90%AD%E5%BB%BA%E4%B8%AA%E4%BA%BA%E5%9B%BE%E5%BA%8A/","title":"用 Cloudflare R2 搭建个人图床：上传、压缩、访问与成本","summary":"从只用 R2 控制台上传，到基于 Cloudflare Workers、D1、Pages 和 TinyPNG 搭建一个可用的个人图床应用。","content_html":"<p>我最早用 Cloudflare R2 做图床时，只是把图片丢到 R2 控制台，再自己拼 URL。这个方式能用，但不适合长期写博客：上传麻烦、图片不好找、无法一键复制地址，也没有压缩流程。</p>\n<p>后来补了一版可视化图床，又加了 TinyPNG 自动压缩。现在回头看，这三篇内容其实应该合成一篇完整方案：R2 负责存储，Workers 负责上传/查询/删除，D1 记录图片元数据，Pages 托管前端页面，TinyPNG 在上传前压缩图片。</p>\n<p><a href=\"/en/posts/2023/cloudflare-r2-image-hosting/\">English version: Build a Personal Image Hosting Service with Cloudflare R2</a></p>\n<p>本文记录的是个人博客图床的完整方案。它不追求做成公开 SaaS，只解决个人写作时的几个核心需求：</p>\n<ul>\n<li>上传图片。</li>\n<li>自动压缩图片。</li>\n<li>生成稳定可访问的图片 URL。</li>\n<li>查看图片列表。</li>\n<li>复制 Markdown 图片地址。</li>\n<li>删除不再需要的图片。</li>\n<li>尽量少花钱，最好在个人用量下接近免费。</li>\n</ul>\n<h2>为什么选 R2</h2>\n<p>个人博客如果是静态生成，正文用 Markdown 管理很舒服，但图片会变成一个麻烦点。</p>\n<p>图片放在博客仓库里，优点是简单，缺点是仓库越来越大，迁移和构建都不舒服。图片放在云服务器上，也能用，但个人服务器带宽通常很小，不值得把图片流量压到服务器上。</p>\n<p>对象存储更适合做图床。Cloudflare R2 的好处是：</p>\n<ul>\n<li>对象存储模型简单，适合存图片。</li>\n<li>可以绑定自定义域名。</li>\n<li>不收取出口流量费，个人博客这类读多写少场景很友好。</li>\n<li>可以和 Workers、D1、Pages 放在同一个平台里组合使用。</li>\n</ul>\n<p>Cloudflare 的价格和免费额度以官方页面为准，本文只讨论适合个人图床的计费结构和使用取舍：<a href=\"https://developers.cloudflare.com/r2/pricing/\">Cloudflare R2 Pricing</a>。</p>\n<p>R2 不是唯一选择。阿里云 OSS、腾讯云 COS、AWS S3 都能做类似事情。选 R2 主要是因为个人博客访问以静态图片为主，R2 的出口流量策略和 Cloudflare 生态比较适合这个场景。</p>\n<h2>整体架构</h2>\n<p>最终架构是这样：</p>\n<p><img src=\"https://aipaint.lihuanyu.com/2023-12-04/cloudflare%E5%9B%BE%E5%BA%8A%E6%9E%B6%E6%9E%84.jpg\" alt=\"cloudflare个人图床架构\"></p>\n<p>涉及的 Cloudflare 服务：</p>\n<ul>\n<li>R2：存储图片文件。</li>\n<li>D1：存储图片元数据，比如文件名、访问地址、创建时间、大小。</li>\n<li>Workers：提供上传、查询、删除 API。</li>\n<li>Pages：托管前端页面。</li>\n</ul>\n<p>额外服务：</p>\n<ul>\n<li>GitHub：保存前端和 Worker 代码。</li>\n<li>TinyPNG/Tinify：压缩图片。</li>\n<li>自定义域名：提供可长期使用的图片访问地址。</li>\n</ul>\n<p>最终成品大概是这样：</p>\n<p><img src=\"https://aipaint.lihuanyu.com/2023-12-04/%E5%9B%BE%E5%BA%8A%E5%BA%94%E7%94%A8-%E5%88%97%E8%A1%A8.png\" alt=\"图床应用图片列表\"></p>\n<p><img src=\"https://aipaint.lihuanyu.com/2023-12-04/%E5%9B%BE%E5%BA%8A%E5%BA%94%E7%94%A8-%E4%B8%8A%E4%BC%A0%E5%9B%BE%E7%89%87.jpg\" alt=\"图床应用上传图片\"></p>\n<h2>准备 R2</h2>\n<p>在 Cloudflare 控制台创建一个 R2 bucket，例如：</p>\n<pre><code class=\"language-text\">image-storage\n</code></pre>\n<p>创建完成后，先解决公开访问问题。R2 默认不公开，图床需要能通过 URL 访问图片。</p>\n<p>Cloudflare 提供两种方式：</p>\n<ul>\n<li>使用 R2 的公开访问能力。</li>\n<li>绑定自己的自定义域名。</li>\n</ul>\n<p>个人博客更适合绑定自己的域名，比如：</p>\n<pre><code class=\"language-text\">https://aipaint.lihuanyu.com\n</code></pre>\n<p>Cloudflare 关于公开桶和自定义域名的说明见：<a href=\"https://developers.cloudflare.com/r2/buckets/public-buckets/\">Public buckets and custom domains</a>。</p>\n<p>长期依赖 Cloudflare 分配的 <code>r2.dev</code> 预览域名风险较高。一方面它不适合正式生产使用，另一方面国内访问也不一定稳定。自己的域名更适合放进历史文章里长期使用。</p>\n<h2>准备 D1</h2>\n<p>R2 只负责存对象，不适合承担图片列表、搜索、分页等功能。这里需要一张表记录图片元数据。</p>\n<p>创建 D1 数据库，例如：</p>\n<pre><code class=\"language-text\">image-storage-record\n</code></pre>\n<p>建表 SQL：</p>\n<pre><code class=\"language-sql\">CREATE TABLE IF NOT EXISTS images (\n  id INTEGER PRIMARY KEY AUTOINCREMENT,\n  object_key TEXT NOT NULL UNIQUE,\n  original_name TEXT NOT NULL,\n  image_url TEXT NOT NULL,\n  content_type TEXT,\n  size INTEGER NOT NULL DEFAULT 0,\n  created_at INTEGER NOT NULL\n);\n\nCREATE INDEX IF NOT EXISTS idx_images_created_at ON images(created_at);\n</code></pre>\n<p>字段含义：</p>\n<ul>\n<li><code>object_key</code>：R2 里的对象 key，例如 <code>2026-05-03/uuid.png</code>。</li>\n<li><code>original_name</code>：用户上传时的原始文件名。</li>\n<li><code>image_url</code>：公开访问地址。</li>\n<li><code>content_type</code>：图片 MIME 类型。</li>\n<li><code>size</code>：最终写入 R2 的文件大小。</li>\n<li><code>created_at</code>：创建时间戳。</li>\n</ul>\n<p>个人场景下，D1 足够承担这类元数据存储。引入 Postgres 或 MySQL 反而会把简单问题复杂化。</p>\n<h2>Worker 绑定配置</h2>\n<p>Worker 通过 binding 访问 R2 和 D1。<code>wrangler.toml</code> 可以这样写：</p>\n<pre><code class=\"language-toml\">name = &quot;image-storage-worker&quot;\nmain = &quot;src/index.ts&quot;\ncompatibility_date = &quot;2026-05-03&quot;\n\n[vars]\nPUBLIC_IMAGE_BASE_URL = &quot;https://aipaint.lihuanyu.com&quot;\n\n[[r2_buckets]]\nbinding = &quot;IMAGE_BUCKET&quot;\nbucket_name = &quot;image-storage&quot;\n\n[[d1_databases]]\nbinding = &quot;DB&quot;\ndatabase_name = &quot;image-storage-record&quot;\ndatabase_id = &quot;替换为 D1 database id&quot;\n</code></pre>\n<p>这里容易漏掉的是 <code>database_id</code>。用 <code>wrangler d1 create image-storage-record</code> 创建数据库时，命令行会返回这个值；如果是在控制台创建，也可以在数据库详情里找到。</p>\n<p>如果要接入 TinyPNG，API key 应通过 secret 管理，而不是写进 <code>wrangler.toml</code>：</p>\n<pre><code class=\"language-bash\">wrangler secret put TINIFY_API_KEY\n</code></pre>\n<p>如果图床页面不希望公开上传，还应增加一个管理 token：</p>\n<pre><code class=\"language-bash\">wrangler secret put ADMIN_TOKEN\n</code></pre>\n<p>Worker binding 的完整配置方式见：<a href=\"https://developers.cloudflare.com/workers/wrangler/configuration/\">Wrangler configuration</a>。</p>\n<h2>Worker API</h2>\n<p>下面是一份简化但完整的 Worker 逻辑，包含：</p>\n<ul>\n<li><code>OPTIONS</code>：处理 CORS 预检。</li>\n<li><code>POST /upload</code>：上传图片，支持 TinyPNG 压缩。</li>\n<li><code>GET /query</code>：分页查询图片列表。</li>\n<li><code>DELETE /delete?id=1</code>：删除图片和元数据。</li>\n</ul>\n<pre><code class=\"language-ts\">interface Env {\n  IMAGE_BUCKET: R2Bucket;\n  DB: D1Database;\n  PUBLIC_IMAGE_BASE_URL: string;\n  TINIFY_API_KEY?: string;\n  ADMIN_TOKEN?: string;\n}\n\nconst corsHeaders = {\n  'Access-Control-Allow-Origin': '*',\n  'Access-Control-Allow-Methods': 'GET,POST,DELETE,OPTIONS',\n  'Access-Control-Allow-Headers': 'Content-Type,Authorization',\n};\n\nexport default {\n  async fetch(request: Request, env: Env): Promise&lt;Response&gt; {\n    const url = new URL(request.url);\n\n    if (request.method === 'OPTIONS') {\n      return new Response(null, { headers: corsHeaders });\n    }\n\n    if (request.method === 'GET' &amp;&amp; url.pathname === '/query') {\n      return handleQuery(request, env);\n    }\n\n    if (!isAuthorized(request, env)) {\n      return json({ success: false, message: 'Unauthorized' }, 401);\n    }\n\n    if (request.method === 'POST' &amp;&amp; url.pathname === '/upload') {\n      return handleUpload(request, env);\n    }\n\n    if (request.method === 'DELETE' &amp;&amp; url.pathname === '/delete') {\n      return handleDelete(request, env);\n    }\n\n    return json({ success: false, message: 'Not found' }, 404);\n  },\n};\n\nfunction isAuthorized(request: Request, env: Env) {\n  if (!env.ADMIN_TOKEN) {\n    return true;\n  }\n\n  return request.headers.get('Authorization') === `Bearer ${env.ADMIN_TOKEN}`;\n}\n\nasync function handleUpload(request: Request, env: Env) {\n  const formData = await request.formData();\n  const file = formData.get('file');\n\n  if (!(file instanceof File)) {\n    return json({ success: false, message: 'Missing file' }, 400);\n  }\n\n  if (!file.type.startsWith('image/')) {\n    return json({ success: false, message: 'Only image files are allowed' }, 400);\n  }\n\n  const objectKey = createObjectKey(file.name);\n  const image = env.TINIFY_API_KEY\n    ? await compressWithTinify(file, env.TINIFY_API_KEY)\n    : {\n        body: await file.arrayBuffer(),\n        contentType: file.type || 'application/octet-stream',\n        size: file.size,\n      };\n\n  await env.IMAGE_BUCKET.put(objectKey, image.body, {\n    httpMetadata: {\n      contentType: image.contentType,\n    },\n  });\n\n  const baseUrl = env.PUBLIC_IMAGE_BASE_URL.replace(/\\/$/, '');\n  const imageUrl = `${baseUrl}/${objectKey}`;\n  const createdAt = Date.now();\n\n  await env.DB.prepare(\n    `INSERT INTO images\n      (object_key, original_name, image_url, content_type, size, created_at)\n     VALUES (?, ?, ?, ?, ?, ?)`,\n  )\n    .bind(objectKey, file.name, imageUrl, image.contentType, image.size, createdAt)\n    .run();\n\n  return json({\n    success: true,\n    url: imageUrl,\n    markdown: `![${file.name}](${imageUrl})`,\n  });\n}\n\nasync function handleQuery(request: Request, env: Env) {\n  const url = new URL(request.url);\n  const pageNum = Math.max(Number(url.searchParams.get('pageNum')) || 1, 1);\n  const pageSize = Math.min(Math.max(Number(url.searchParams.get('pageSize')) || 20, 1), 50);\n  const offset = (pageNum - 1) * pageSize;\n\n  const list = await env.DB.prepare(\n    `SELECT id, object_key, original_name, image_url, content_type, size, created_at\n     FROM images\n     ORDER BY id DESC\n     LIMIT ? OFFSET ?`,\n  )\n    .bind(pageSize, offset)\n    .all();\n\n  const count = await env.DB.prepare(`SELECT COUNT(*) AS total FROM images`).first&lt;{\n    total: number;\n  }&gt;();\n\n  return json({\n    success: true,\n    results: list.results,\n    total: count?.total || 0,\n  });\n}\n\nasync function handleDelete(request: Request, env: Env) {\n  const url = new URL(request.url);\n  const id = Number(url.searchParams.get('id'));\n\n  if (!Number.isInteger(id) || id &lt;= 0) {\n    return json({ success: false, message: 'Invalid id' }, 400);\n  }\n\n  const row = await env.DB.prepare(`SELECT object_key FROM images WHERE id = ?`)\n    .bind(id)\n    .first&lt;{ object_key: string }&gt;();\n\n  if (!row) {\n    return json({ success: false, message: 'Image not found' }, 404);\n  }\n\n  await env.IMAGE_BUCKET.delete(row.object_key);\n  await env.DB.prepare(`DELETE FROM images WHERE id = ?`).bind(id).run();\n\n  return json({ success: true });\n}\n\nasync function compressWithTinify(file: File, apiKey: string) {\n  const source = await file.arrayBuffer();\n  const auth = `Basic ${btoa(`api:${apiKey}`)}`;\n\n  const shrink = await fetch('https://api.tinify.com/shrink', {\n    method: 'POST',\n    headers: {\n      Authorization: auth,\n      'Content-Type': file.type || 'application/octet-stream',\n    },\n    body: source,\n  });\n\n  if (!shrink.ok) {\n    const message = await shrink.text();\n    throw new Error(`TinyPNG shrink failed: ${shrink.status} ${message}`);\n  }\n\n  const outputUrl = shrink.headers.get('Location');\n\n  if (!outputUrl) {\n    throw new Error('TinyPNG did not return output location');\n  }\n\n  const optimized = await fetch(outputUrl, {\n    headers: {\n      Authorization: auth,\n    },\n  });\n\n  if (!optimized.ok) {\n    const message = await optimized.text();\n    throw new Error(`TinyPNG download failed: ${optimized.status} ${message}`);\n  }\n\n  const body = await optimized.arrayBuffer();\n\n  return {\n    body,\n    contentType: optimized.headers.get('Content-Type') || file.type || 'application/octet-stream',\n    size: Number(optimized.headers.get('Content-Length')) || body.byteLength,\n  };\n}\n\nfunction createObjectKey(filename: string) {\n  const extension = filename.includes('.') ? filename.split('.').pop() : 'bin';\n  const date = new Date().toISOString().slice(0, 10);\n  return `${date}/${crypto.randomUUID()}.${extension}`;\n}\n\nfunction json(data: unknown, status = 200) {\n  return new Response(JSON.stringify(data), {\n    status,\n    headers: {\n      ...corsHeaders,\n      'Content-Type': 'application/json; charset=utf-8',\n    },\n  });\n}\n</code></pre>\n<p>这里有几个细节值得强调。</p>\n<p>第一，上传接口要校验 <code>image/*</code>，否则图床很容易变成任意文件存储。</p>\n<p>第二，个人使用也应加 <code>ADMIN_TOKEN</code>。前端请求时带上：</p>\n<pre><code class=\"language-text\">Authorization: Bearer 管理 token\n</code></pre>\n<p>第三，TinyPNG 的 API 不是从 JSON 里拿 <code>output.url</code>。压缩请求成功后，应从响应头的 <code>Location</code> 获取压缩结果地址，再请求这个地址下载压缩后的图片。接口细节见：<a href=\"https://tinypng.com/developers/reference\">Tinify API reference</a>。</p>\n<p>第四，<code>object_key</code> 直接使用原始文件名会带来中文、空格、同名覆盖等问题。按日期加 UUID 更稳。</p>\n<h2>前端页面</h2>\n<p>前端可以用任何框架。示例版本使用的是 SolidJS，换成 React、Vue、Svelte 也没有本质差异，图床不是复杂应用。</p>\n<p>核心功能其实只有三个。</p>\n<p>上传：</p>\n<pre><code class=\"language-ts\">async function uploadImage(file: File) {\n  const formData = new FormData();\n  formData.append('file', file);\n\n  const response = await fetch(`${apiBaseUrl}/upload`, {\n    method: 'POST',\n    headers: {\n      Authorization: `Bearer ${adminToken}`,\n    },\n    body: formData,\n  });\n\n  if (!response.ok) {\n    throw new Error(await response.text());\n  }\n\n  return response.json();\n}\n</code></pre>\n<p>查询：</p>\n<pre><code class=\"language-ts\">async function queryImages(pageNum = 1, pageSize = 20) {\n  const response = await fetch(\n    `${apiBaseUrl}/query?pageNum=${pageNum}&amp;pageSize=${pageSize}`,\n  );\n\n  if (!response.ok) {\n    throw new Error(await response.text());\n  }\n\n  return response.json();\n}\n</code></pre>\n<p>删除：</p>\n<pre><code class=\"language-ts\">async function deleteImage(id: number) {\n  const response = await fetch(`${apiBaseUrl}/delete?id=${id}`, {\n    method: 'DELETE',\n    headers: {\n      Authorization: `Bearer ${adminToken}`,\n    },\n  });\n\n  if (!response.ok) {\n    throw new Error(await response.text());\n  }\n\n  return response.json();\n}\n</code></pre>\n<p>页面上至少需要这些交互：</p>\n<ul>\n<li>文件选择或拖拽上传。</li>\n<li>上传成功后展示图片 URL 和 Markdown。</li>\n<li>图片列表展示缩略图、原始文件名、创建时间、大小。</li>\n<li>复制 URL。</li>\n<li>复制 Markdown。</li>\n<li>删除图片。</li>\n</ul>\n<p>前端部署到 Cloudflare Pages 即可。它和 Worker API 可以分开部署，也可以共用一个仓库。个人项目里分开维护更清晰：前端页面出问题不影响图片访问，Worker API 也更容易单独迭代。</p>\n<h2>自定义域名与缓存</h2>\n<p>图床最重要的是 URL 稳定。只要图片 URL 被写进文章，就不应该轻易变化。</p>\n<p>推荐做法：</p>\n<ul>\n<li>单独给图片服务一个子域名，例如 <code>aipaint.lihuanyu.com</code>。</li>\n<li>R2 bucket 绑定这个子域名。</li>\n<li>文章里只使用这个子域名下的图片地址。</li>\n<li>避免把 Worker 预览域名、Pages 预览域名写进文章。</li>\n</ul>\n<p>图片属于静态资源，缓存可以激进一点。个人博客图片一旦上传，通常不会用同一个 URL 替换内容。如果确实要替换，最简单的方式是上传新图片，生成新 URL。</p>\n<h2>成本</h2>\n<p>这套方案的成本主要来自四块：</p>\n<ul>\n<li>R2 存储和请求。</li>\n<li>D1 读写。</li>\n<li>Workers 请求。</li>\n<li>TinyPNG 压缩次数。</li>\n</ul>\n<p>个人博客的图片访问通常是读多写少，R2 和 Workers 的压力都很小。真正需要单独评估的是 TinyPNG：它不是 Cloudflare 服务，免费额度和计费规则要看 Tinify 官方说明。如果图片很多，压缩可以改成本地脚本处理，或者只在上传大图时启用。</p>\n<p>实践后的取舍是：</p>\n<ul>\n<li>R2 适合长期存图片。</li>\n<li>D1 只存元数据，成本可以忽略。</li>\n<li>Workers 很适合做这类轻量 API。</li>\n<li>TinyPNG 是锦上添花，不是必需项。</li>\n</ul>\n<p>如果只是写博客，第一版可以先不做 TinyPNG，把上传、查询、复制 Markdown 跑通。等图片变多、加载速度开始成为问题，再接压缩。</p>\n<h2>这套方案的边界</h2>\n<p>它适合个人图床，不适合直接做公开平台。</p>\n<p>如果要开放给其他人用，至少还要补：</p>\n<ul>\n<li>用户系统。</li>\n<li>权限隔离。</li>\n<li>上传频率限制。</li>\n<li>文件大小限制。</li>\n<li>内容安全审核。</li>\n<li>存储配额。</li>\n<li>删除后的审计日志。</li>\n<li>防盗链或访问控制策略。</li>\n</ul>\n<p>个人使用时，最重要的是不要暴露上传接口。否则上传接口可能被滥用为公开文件存储。</p>\n<h2>结论</h2>\n<p>Cloudflare R2 做个人图床是合适的，但不要停留在“控制台上传 + 手拼 URL”的阶段。真正好用的图床，需要把上传、压缩、列表、复制、删除串起来。</p>\n<p>落地顺序可以是：</p>\n<ol>\n<li>先创建 R2，并绑定自己的图片域名。</li>\n<li>再用 Worker 写上传接口。</li>\n<li>用 D1 记录图片元数据。</li>\n<li>做一个很简单的前端页面。</li>\n<li>最后再接 TinyPNG 或其他压缩方案。</li>\n</ol>\n<p>这样既保留了对象存储的稳定性，又能让写博客时贴图这件事变得足够顺手。</p>\n<h2>参考链接</h2>\n<ul>\n<li><a href=\"https://developers.cloudflare.com/r2/pricing/\">Cloudflare R2 Pricing</a></li>\n<li><a href=\"https://developers.cloudflare.com/r2/buckets/public-buckets/\">Cloudflare R2 Public buckets</a></li>\n<li><a href=\"https://developers.cloudflare.com/workers/wrangler/configuration/\">Cloudflare Workers Wrangler configuration</a></li>\n<li><a href=\"https://tinypng.com/developers/reference\">Tinify API reference</a></li>\n</ul>\n","date_published":"2023-12-04T00:00:00.000Z","date_modified":"2026-05-03T00:00:00.000Z","tags":["图床","Cloudflare","R2","D1","Worker","TinyPNG"],"language":"zh"},{"id":"https://www.lihuanyu.com/posts/2023/%E8%8B%B1%E8%AF%AD%E7%9A%84%E5%8F%A3%E9%9F%B3/","url":"https://www.lihuanyu.com/posts/2023/%E8%8B%B1%E8%AF%AD%E7%9A%84%E5%8F%A3%E9%9F%B3/","title":"英语的口音","summary":"从印度、日本和中国人说英语的口音差异出发，解释清浊、送气、不送气等发音差异为什么会影响英语可理解度。","content_html":"<blockquote>\n<p>一个嘲笑印度、日本人说英语的笑话：两个印度人在嘲笑日本人的英语发。“Jabonese agcent is vedy, vedy hard to undershdand.” 然后被日本人神吐槽 “Indeian ekusento ishi belly belly haudo tsu andasudando.”</p>\n</blockquote>\n<p>以前总觉得中国人说英文的发音还是蛮标准的，至少比印度人、日本人好多了，结果知乎上逛到一个毁三观的结论。对于英国/美国人来说，说英语便于理解的程度大概是，印度人 &gt;&gt; 日本人 ≈ 中国人。</p>\n<p>首先思考一个问题：为什么印度人或日本人说英语，你一听就能听出来？发音差了那么多，难道他们自己注意不到吗？<br>\n对，他们自己注意不到。<br>\n<strong>反之亦然，中国人说英语，其中发音离谱的地方，中国人自己也注意不到。</strong></p>\n<p>我们都觉得印度人“清浊不分”，把 p,t,k 读成 b,d,g 。举例：印度人说 me too 非常像 me do 。</p>\n<p>而事实上，真正清浊不分的是我们自己，印度人 p,t,k 发的就是清音，只不过是不送气清音，相当于 speak 、 star 、 skin 里面的 p,t,k ，但在中国人听起来，是和浊音 b,d,g 没什么区别的（还给这种现象起了个名叫“清音浊化”，但其实那仍然是清音）。</p>\n<p>原因是，汉语不是清浊对立，而是强弱对立，也就是送气清音与不送气清音，汉语普通话里根本不存在真正的 b,d,g 浊音。换句话说，我们（业余英语）在发本该是浊音的 b,d,g 时，发出来的音其实都是不送气的清音 p,t,k 。</p>\n<p>请思考一下北京、豆腐、功夫、宫保鸡丁的英文音译 —— <strong>P</strong>eking、<strong>T</strong>ofu、<strong>K</strong>ungfu、<strong>K</strong>ung <strong>P</strong>ao Chicken —— 明白了吧？我们以为自己发音是 b,d,g ，但实际发出来的根本就是 p,t,k 。只不过是不送气或送气程度较弱的 p,t,k 而已。（这段内容真的解释了我长久以来的困惑，为什么会翻译成这样呢，原来在老外耳里，我们说的就是 p、t、k ）</p>\n<p>如果把塞音进行区分，一分是清（unvoiced）和浊（voiced）；另一分是送气（aspirated）和不送气（unaspirated），即汉语里的强弱区分。<br>\n<strong>清浊的区别是声带是否发声，发音的时候摸摸喉咙就知道。送气和不送气的区别，发音的时候拿一只手放在嘴前面感受一下有没有爆破的气流。</strong></p>\n<p>来现在把手放在喉咙上发音感受下，用标准的普通话说“拜拜”，是不是能感受到喉咙几乎没发音。这时再说“崇拜”，这个拜字就能明显感受到喉咙在发音了。这是少有的能发出浊音的汉字，还得借助连读。平时常用的字音几乎没有浊音的，所以中国人分不清印度人的清浊区分读法，而印度人又不分送气不送气，中国人就觉得他们简直在乱说。</p>\n<p><img src=\"https://aipaint.lihuanyu.com/%E6%B8%85%E6%B5%8A%E9%9F%B3%E5%8C%BA%E5%88%AB.jpg\" alt=\"塞音四分\"></p>\n<p>如图，美国人发音是清浊音伴随着送不送气的区别。印度人发音只关注清浊音，送不送气对他们都是同一个音。中国人发音只关注送不送气（上文提到的强弱对立），清浊反而分不太清。</p>\n<p>中国人靠送气区分t和d，印度人靠声带发声来区别t和d，而大多数情况下美国人说的t和d，既有送气和不送气的分别，又有声带发不发声的区别。所以中国人印度人互相听着费劲，而都容易理解美国人。只有在美国人说stop这个词的时候，才说不送气轻音，中国人听起来就觉得是d了。相比之下，印度人理解中国人比中国人理解印度人还容易一些，因为这四个音他们都有。</p>\n<p>普通话说“大”的时候，实际上音标上是不送气的/ta/。而如果我们说普通话时把“大”发成英语里/da/的音，听起来就像嗓子顿了一下似的，像很多外国人说汉语的那种洋味十足。（个人的细节感受就是发音位置的区别，正常普通话的大，发音位置就在牙齿舌头那个位置的感觉，而如果用/da/的音，发音位置就在喉咙声带那个位置）</p>\n<p>印度人在应该发送气清音的时候，发出的却经常是不送气清音，这是因为他们的语言中根本不存在送气不送气的区别，送气和不送气，在他们概念里是一样的。（就像部分四川人平翘不分、山东人日乐不分、甘肃人云勇不分，是他们不愿意分吗?是真听不出来差别）<br>\n相应的，我们在遇到该发浊音的时候，发出来的一般都是不送气清音，只是我们自己注意不到，因为浊音与不送气清音，在我们的概念里是一样的。</p>\n<p>所以，印度人的发音虽然在中国人听来和正统英语相去甚远，但实际它就是一种英语的方言，更容易被美国人英国人理解。不过现代英式英语的发音，浊音在渐渐弱化，十分接近不送气清音，反而与汉语接近了，所以中国人把bdg念成不送气清音，英国人听起来应该是没什么障碍，而美式英语仍然是较为明显的清浊对立。</p>\n<p>突然想起来最近的封神榜电影里，费翔的商务殷语，大概也就是这么回事。</p>\n<p>还有一个点是，印度当年是英国殖民地，现在英语也是他们的官方语言之一。这意味着民众都能张口说，在频繁的使用过程中，口音是会趋于一致的，就算说得和标准不一样，只要听懂了一个就能听懂所有印度人说话，而中国人的官方语言就是中文，英语只用于考试和阅读，很少用于交流，所以可能会按学校甚至按个人都存在非常大的区别，当然会更难以让英语母语者(native speaker)听明白。</p>\n","date_published":"2023-11-12T00:00:00.000Z","tags":["英语"],"language":"zh"},{"id":"https://www.lihuanyu.com/en/posts/2023/nestjs-interceptors-skip-response-wrapping/","url":"https://www.lihuanyu.com/en/posts/2023/nestjs-interceptors-skip-response-wrapping/","title":"How to Skip a Global NestJS Interceptor for Specific Routes","summary":"Skip a global NestJS response interceptor on selected routes with SetMetadata and Reflector, while keeping response wrapping enabled everywhere else.","content_html":"<p>Use a custom decorator and NestJS <code>Reflector</code> metadata to skip a global interceptor for one route or an entire controller. This pattern works well when most API responses use one JSON wrapper but webhooks, file downloads, streams, or proxy endpoints must return their original response.</p>\n<p><a href=\"/posts/2023/nestjs%E6%8B%A6%E6%88%AA%E5%99%A8%E4%B8%8E%E8%B7%B3%E8%BF%87%E6%8B%A6%E6%88%AA%E5%99%A8/\">Chinese version of this article</a></p>\n<h2>Mark routes that skip the interceptor</h2>\n<p>Define a metadata key and expose it through a decorator:</p>\n<pre><code class=\"language-ts\">import { SetMetadata } from '@nestjs/common';\n\nexport const SKIP_RESPONSE_WRAP = 'skipResponseWrap';\n\nexport const SkipResponseWrap = () =&gt;\n  SetMetadata(SKIP_RESPONSE_WRAP, true);\n</code></pre>\n<p>Apply <code>@SkipResponseWrap()</code> to a route that must return its original value:</p>\n<pre><code class=\"language-ts\">import { Body, Controller, HttpCode, Post } from '@nestjs/common';\nimport { SkipResponseWrap } from './skip-response-wrap.decorator';\n\n@Controller('webhooks')\nexport class WebhookController {\n  @Post('provider')\n  @HttpCode(200)\n  @SkipResponseWrap()\n  handleWebhook(@Body() body: unknown) {\n    return 'success';\n  }\n}\n</code></pre>\n<p>Place the decorator on the controller class when every route in that controller should skip wrapping.</p>\n<h2>Read skip metadata in the global interceptor</h2>\n<p>Start with the interceptor imports and response type:</p>\n<pre><code class=\"language-ts\">import {\n  CallHandler,\n  ExecutionContext,\n  Injectable,\n  NestInterceptor,\n} from '@nestjs/common';\nimport { Reflector } from '@nestjs/core';\nimport { map } from 'rxjs/operators';\nimport { SKIP_RESPONSE_WRAP } from './skip-response-wrap.decorator';\n\ninterface ApiResponse&lt;T&gt; {\n  code: number;\n  success: true;\n  data: T;\n}\n</code></pre>\n<p>Then inject <code>Reflector</code> and check metadata before applying the response mapping:</p>\n<pre><code class=\"language-ts\">@Injectable()\nexport class ResponseWrapInterceptor&lt;T&gt;\n  implements NestInterceptor&lt;T, ApiResponse&lt;T&gt; | T&gt;\n{\n  constructor(private readonly reflector: Reflector) {}\n\n  intercept(context: ExecutionContext, next: CallHandler&lt;T&gt;) {\n    const skip = this.reflector.getAllAndOverride&lt;boolean&gt;(\n      SKIP_RESPONSE_WRAP,\n      [context.getHandler(), context.getClass()],\n    );\n\n    if (skip) return next.handle();\n\n    return next.handle().pipe(\n      map((data) =&gt; ({ code: 0, success: true, data })),\n    );\n  }\n}\n</code></pre>\n<p><code>context.getHandler()</code> reads route-level metadata. <code>context.getClass()</code> reads controller-level metadata. <code>getAllAndOverride()</code> lets the route setting take precedence over the controller setting.</p>\n<h2>Register the interceptor with dependency injection</h2>\n<p>Register the interceptor through <code>APP_INTERCEPTOR</code> so NestJS can inject <code>Reflector</code>:</p>\n<pre><code class=\"language-ts\">import { Module } from '@nestjs/common';\nimport { APP_INTERCEPTOR } from '@nestjs/core';\nimport { ResponseWrapInterceptor } from './response-wrap.interceptor';\n\n@Module({\n  providers: [\n    {\n      provide: APP_INTERCEPTOR,\n      useClass: ResponseWrapInterceptor,\n    },\n  ],\n})\nexport class AppModule {}\n</code></pre>\n<p>Most controllers can now return business data directly. The interceptor converts successful results into a shared shape:</p>\n<pre><code class=\"language-json\">{\n  &quot;code&quot;: 0,\n  &quot;success&quot;: true,\n  &quot;data&quot;: {\n    &quot;user&quot;: &quot;example_user&quot;,\n    &quot;imageUrl&quot;: &quot;https://example.com/avatar.png&quot;\n  }\n}\n</code></pre>\n<p>The decorated webhook route returns its required plain-text value instead:</p>\n<pre><code class=\"language-text\">success\n</code></pre>\n<h2>Keep error responses in exception filters</h2>\n<p>Some implementations use another interceptor with <code>catchError()</code> to wrap errors. That can work, but it is not always the clearest boundary.</p>\n<p>Wrapping successful responses is a transformation after a handler returns normally, which fits an interceptor well. Error responses are exception handling, and NestJS has exception filters for that. Keeping the two paths separate reduces the chance of swallowing exceptions inside a response interceptor.</p>\n<p>A simplified HTTP exception filter can look like this:</p>\n<pre><code class=\"language-ts\">import {\n  ArgumentsHost,\n  Catch,\n  ExceptionFilter,\n  HttpException,\n  HttpStatus,\n} from '@nestjs/common';\nimport { Response } from 'express';\n\n@Catch()\nexport class HttpExceptionFilter implements ExceptionFilter {\n  catch(exception: unknown, host: ArgumentsHost) {\n    const ctx = host.switchToHttp();\n    const response = ctx.getResponse&lt;Response&gt;();\n\n    const status =\n      exception instanceof HttpException\n        ? exception.getStatus()\n        : HttpStatus.INTERNAL_SERVER_ERROR;\n\n    const message =\n      exception instanceof HttpException\n        ? exception.message\n        : 'Internal server error';\n\n    response.status(status).json({\n      code: status,\n      success: false,\n      message,\n    });\n  }\n}\n</code></pre>\n<p>Register it globally in <code>main.ts</code>:</p>\n<pre><code class=\"language-ts\">const app = await NestFactory.create(AppModule);\napp.useGlobalFilters(new HttpExceptionFilter());\nawait app.listen(3000);\n</code></pre>\n<p>Real projects can add business error codes, request IDs, logging, and custom exception classes. The core boundary stays the same: use an interceptor for successful response mapping, and use a filter for exception responses.</p>\n<h2>Decide which endpoints should skip wrapping</h2>\n<p>A global interceptor affects every endpoint. Some endpoints should not return a JSON wrapper:</p>\n<ul>\n<li>Webhooks from WeChat, GitHub, Stripe, and similar platforms may require a fixed body or status code.</li>\n<li>File download endpoints need to return streams.</li>\n<li>Images, QR codes, and CSV exports have their own <code>Content-Type</code>.</li>\n<li>Proxy endpoints may need to pass through the upstream response.</li>\n</ul>\n<p>If the interceptor wraps one of these responses as <code>{ code, success, data }</code>, the caller may reject it. Mark that endpoint explicitly instead of adding response-type guesses to the interceptor.</p>\n<h2>Account for response mapping boundaries</h2>\n<p>The NestJS documentation includes an important warning: response mapping does not work with the library-specific response strategy, such as directly using the <code>@Res()</code> object in a handler.</p>\n<p>A practical split is:</p>\n<ul>\n<li>Normal JSON APIs: return business objects and let the global interceptor wrap them.</li>\n<li>Protocol-specific responses: add <code>@SkipResponseWrap()</code> and return the exact value required.</li>\n<li>File streams or strongly controlled headers: add <code>@SkipResponseWrap()</code> and use <code>@Res()</code> or <code>StreamableFile</code> when needed.</li>\n<li>Error responses: prefer exception filters instead of mixing them into the successful response interceptor.</li>\n</ul>\n<p>This keeps the global rule simple while still giving special endpoints an explicit escape hatch. The interceptor wraps successful responses, the decorator declares exceptions, and the exception filter handles error shape.</p>\n<h2>Further reading</h2>\n<ul>\n<li><a href=\"https://docs.nestjs.com/interceptors\">NestJS Docs: Interceptors</a></li>\n<li><a href=\"https://docs.nestjs.com/fundamentals/execution-context\">NestJS Docs: Execution context, reflection and metadata</a></li>\n</ul>\n","date_published":"2023-10-22T00:00:00.000Z","date_modified":"2026-08-01T00:00:00.000Z","tags":["NestJS","Interceptors","Backend"],"language":"en"},{"id":"https://www.lihuanyu.com/posts/2023/nestjs%E6%8B%A6%E6%88%AA%E5%99%A8%E4%B8%8E%E8%B7%B3%E8%BF%87%E6%8B%A6%E6%88%AA%E5%99%A8/","url":"https://www.lihuanyu.com/posts/2023/nestjs%E6%8B%A6%E6%88%AA%E5%99%A8%E4%B8%8E%E8%B7%B3%E8%BF%87%E6%8B%A6%E6%88%AA%E5%99%A8/","title":"NestJS 拦截器与跳过拦截器","summary":"用 NestJS 全局拦截器统一包装成功响应，并通过自定义装饰器为 Webhook、文件下载等接口跳过包装。","content_html":"<p>写 API 接口时，常见需求是让成功响应保持统一结构。比如业务方法返回用户信息：</p>\n<pre><code class=\"language-json\">{\n  &quot;user&quot;: &quot;xxx&quot;,\n  &quot;imageUrl&quot;: &quot;https://example.com/avatar.png&quot;\n}\n</code></pre>\n<p>接口返回时，希望统一包一层：</p>\n<pre><code class=\"language-json\">{\n  &quot;code&quot;: 0,\n  &quot;success&quot;: true,\n  &quot;data&quot;: {\n    &quot;user&quot;: &quot;xxx&quot;,\n    &quot;imageUrl&quot;: &quot;https://example.com/avatar.png&quot;\n  }\n}\n</code></pre>\n<p>这类逻辑很适合放在 NestJS interceptor 里。拦截器可以在 handler 执行前后介入，也可以用 RxJS operator 改写 handler 的返回值。</p>\n<p><a href=\"/en/posts/2023/nestjs-interceptors-skip-response-wrapping/\">English version: NestJS Interceptors and How to Skip Response Wrapping</a></p>\n<h2>统一成功响应</h2>\n<p>一个最基础的成功响应包装拦截器可以这样写：</p>\n<pre><code class=\"language-ts\">import {\n  CallHandler,\n  ExecutionContext,\n  Injectable,\n  NestInterceptor,\n} from '@nestjs/common';\nimport { Observable } from 'rxjs';\nimport { map } from 'rxjs/operators';\n\ninterface ApiResponse&lt;T&gt; {\n  code: number;\n  success: true;\n  data: T;\n}\n\n@Injectable()\nexport class ResponseWrapInterceptor&lt;T&gt;\n  implements NestInterceptor&lt;T, ApiResponse&lt;T&gt;&gt;\n{\n  intercept(\n    context: ExecutionContext,\n    next: CallHandler&lt;T&gt;,\n  ): Observable&lt;ApiResponse&lt;T&gt;&gt; {\n    return next.handle().pipe(\n      map((data) =&gt; ({\n        code: 0,\n        success: true,\n        data,\n      })),\n    );\n  }\n}\n</code></pre>\n<p>这里的关键点是 <code>next.handle()</code> 返回的是一个 <code>Observable</code>。handler 原本返回的数据会进入这个流，<code>map()</code> 可以把它转换成统一结构。</p>\n<p>注册成全局拦截器时，推荐使用 <code>APP_INTERCEPTOR</code>，这样拦截器仍然在 Nest 的依赖注入上下文里：</p>\n<pre><code class=\"language-ts\">import { Module } from '@nestjs/common';\nimport { APP_INTERCEPTOR } from '@nestjs/core';\nimport { ResponseWrapInterceptor } from './response-wrap.interceptor';\n\n@Module({\n  providers: [\n    {\n      provide: APP_INTERCEPTOR,\n      useClass: ResponseWrapInterceptor,\n    },\n  ],\n})\nexport class AppModule {}\n</code></pre>\n<p>这样大多数接口只需要返回业务数据，不需要每个 controller 都手写 <code>{ code, success, data }</code>。</p>\n<h2>错误响应更适合放在异常过滤器</h2>\n<p>有些代码会用另一个 interceptor 配合 <code>catchError()</code> 把异常也包装起来。这个做法能工作，但不一定是更清晰的边界。</p>\n<p>成功响应包装属于“handler 正常返回后的数据转换”，很适合 interceptor。错误响应则是异常处理，NestJS 里更直接的工具是 exception filter。这样成功和失败两条链路更清楚，也不容易在 interceptor 里误吞异常。</p>\n<p>一个简化的 HTTP 异常过滤器可以这样写：</p>\n<pre><code class=\"language-ts\">import {\n  ArgumentsHost,\n  Catch,\n  ExceptionFilter,\n  HttpException,\n  HttpStatus,\n} from '@nestjs/common';\nimport { Response } from 'express';\n\n@Catch()\nexport class HttpExceptionFilter implements ExceptionFilter {\n  catch(exception: unknown, host: ArgumentsHost) {\n    const ctx = host.switchToHttp();\n    const response = ctx.getResponse&lt;Response&gt;();\n\n    const status =\n      exception instanceof HttpException\n        ? exception.getStatus()\n        : HttpStatus.INTERNAL_SERVER_ERROR;\n\n    const message =\n      exception instanceof HttpException\n        ? exception.message\n        : 'Internal server error';\n\n    response.status(status).json({\n      code: status,\n      success: false,\n      message,\n    });\n  }\n}\n</code></pre>\n<p>全局注册可以放在 <code>main.ts</code>：</p>\n<pre><code class=\"language-ts\">const app = await NestFactory.create(AppModule);\napp.useGlobalFilters(new HttpExceptionFilter());\nawait app.listen(3000);\n</code></pre>\n<p>实际项目里还可以继续补充错误码映射、日志记录、请求 ID、业务异常基类等内容。核心原则是：成功响应用 interceptor 转换，异常响应用 filter 处理。</p>\n<h2>为什么需要跳过包装</h2>\n<p>全局拦截器的问题在于它会影响所有接口。但有些接口不能返回统一 JSON：</p>\n<ul>\n<li>微信、GitHub、Stripe 等 Webhook 可能要求返回固定文本或固定状态码。</li>\n<li>文件下载接口需要直接返回文件流。</li>\n<li>图片、二维码、CSV 导出等接口有自己的 <code>Content-Type</code>。</li>\n<li>代理接口可能需要原样透传上游响应。</li>\n</ul>\n<p>如果这些接口也被包成 <code>{ code, success, data }</code>，调用方就无法按协议识别响应。解决办法是给接口打一个“跳过包装”的标记，让拦截器读这个 metadata。</p>\n<h2>定义跳过包装装饰器</h2>\n<p>可以用 <code>SetMetadata</code> 定义一个装饰器：</p>\n<pre><code class=\"language-ts\">import { SetMetadata } from '@nestjs/common';\n\nexport const SKIP_RESPONSE_WRAP = 'skipResponseWrap';\n\nexport const SkipResponseWrap = () =&gt; SetMetadata(SKIP_RESPONSE_WRAP, true);\n</code></pre>\n<p>然后在拦截器里注入 <code>Reflector</code>，同时读取 handler 和 controller 上的 metadata：</p>\n<pre><code class=\"language-ts\">import {\n  CallHandler,\n  ExecutionContext,\n  Injectable,\n  NestInterceptor,\n} from '@nestjs/common';\nimport { Reflector } from '@nestjs/core';\nimport { Observable } from 'rxjs';\nimport { map } from 'rxjs/operators';\nimport { SKIP_RESPONSE_WRAP } from './skip-response-wrap.decorator';\n\ninterface ApiResponse&lt;T&gt; {\n  code: number;\n  success: true;\n  data: T;\n}\n\n@Injectable()\nexport class ResponseWrapInterceptor&lt;T&gt;\n  implements NestInterceptor&lt;T, ApiResponse&lt;T&gt; | T&gt;\n{\n  constructor(private readonly reflector: Reflector) {}\n\n  intercept(\n    context: ExecutionContext,\n    next: CallHandler&lt;T&gt;,\n  ): Observable&lt;ApiResponse&lt;T&gt; | T&gt; {\n    const skip = this.reflector.getAllAndOverride&lt;boolean&gt;(\n      SKIP_RESPONSE_WRAP,\n      [context.getHandler(), context.getClass()],\n    );\n\n    if (skip) {\n      return next.handle();\n    }\n\n    return next.handle().pipe(\n      map((data) =&gt; ({\n        code: 0,\n        success: true,\n        data,\n      })),\n    );\n  }\n}\n</code></pre>\n<p><code>context.getHandler()</code> 对应当前路由方法，<code>context.getClass()</code> 对应 controller 类。用 <code>getAllAndOverride()</code> 的好处是装饰器既可以放在方法上，也可以放在整个 controller 上。</p>\n<h2>使用方式</h2>\n<p>比如微信消息推送要求服务端返回纯文本 <code>success</code>，就可以给这个接口加上 <code>@SkipResponseWrap()</code>：</p>\n<pre><code class=\"language-ts\">import { Body, Controller, HttpCode, Post } from '@nestjs/common';\nimport { SkipResponseWrap } from './skip-response-wrap.decorator';\n\n@Controller('wechat')\nexport class WechatController {\n  @Post('push')\n  @HttpCode(200)\n  @SkipResponseWrap()\n  async handleWechatPush(@Body() data: unknown) {\n    // 校验签名、处理消息、记录日志等\n    return 'success';\n  }\n}\n</code></pre>\n<p>这样这个接口会返回纯字符串 <code>success</code>，不会被包装成：</p>\n<pre><code class=\"language-json\">{\n  &quot;code&quot;: 0,\n  &quot;success&quot;: true,\n  &quot;data&quot;: &quot;success&quot;\n}\n</code></pre>\n<p>如果一个 controller 下所有接口都不需要包装，也可以把装饰器放在类上：</p>\n<pre><code class=\"language-ts\">@SkipResponseWrap()\n@Controller('files')\nexport class FilesController {}\n</code></pre>\n<h2>需要注意的边界</h2>\n<p>NestJS 文档里有一个重要提醒：response mapping 不适用于直接使用 library-specific response strategy 的场景，也就是在 handler 里直接使用 <code>@Res()</code> 操作原始响应对象。</p>\n<p>所以实践里可以按下面的规则分工：</p>\n<ul>\n<li>普通 JSON API：直接返回业务对象，让全局 interceptor 包装。</li>\n<li>固定协议响应：加 <code>@SkipResponseWrap()</code>，直接返回协议需要的内容。</li>\n<li>文件流或强控制响应头：加 <code>@SkipResponseWrap()</code>，必要时使用 <code>@Res()</code> 或 <code>StreamableFile</code>。</li>\n<li>错误响应：优先交给 exception filter，而不是在成功响应 interceptor 里混着处理。</li>\n</ul>\n<p>这样做的好处是全局规则仍然简单，但特殊接口有明确出口。拦截器负责统一成功响应，装饰器负责声明例外，异常过滤器负责错误结构，三者的职责边界比较清楚。</p>\n<h2>扩展阅读</h2>\n<ul>\n<li><a href=\"https://docs.nestjs.com/interceptors\">NestJS Docs: Interceptors</a></li>\n<li><a href=\"https://docs.nestjs.com/fundamentals/execution-context\">NestJS Docs: Execution context, reflection and metadata</a></li>\n</ul>\n","date_published":"2023-10-22T00:00:00.000Z","date_modified":"2026-05-05T00:00:00.000Z","tags":["nestjs","拦截器","开发"],"language":"zh"},{"id":"https://www.lihuanyu.com/posts/2023/%E4%BD%A0%E5%A5%BD%E5%85%B0%E5%B7%9E/","url":"https://www.lihuanyu.com/posts/2023/%E4%BD%A0%E5%A5%BD%E5%85%B0%E5%B7%9E/","title":"你好兰州","summary":"记录一次去兰州参加婚礼的短途旅行，从硬卧火车、黄河、牛肉面、博物馆到西北婚礼和城市气质。","content_html":"<p>中秋的前一天，关系要好的大学室友要结婚了，请假两天去兰州参加婚礼，草草游览了兰州。</p>\n<p>因为时间和价格的原因，这趟行程往返都是火车，还是硬卧，本来以为睡一觉就能到达，会比飞机更舒服。但实际硬卧车厢的体验并不好，人很多环境比较脏，还有烟味和小孩的吵闹，睡眠质量向来还可以的我都失眠到两三点才迷迷糊糊睡了过去。</p>\n<p>想想自己差不多10多年没坐过硬卧火车了，记忆中的硬卧车厢还是蛮高级的存在，现在却变得如此糟糕。</p>\n<p><img src=\"https://aipaint.lihuanyu.com/%E7%A1%AC%E5%8D%A7%E8%BD%A6%E5%8E%A2.jpeg\" alt=\"硬卧车厢\"></p>\n<p>一觉醒来差不多就在兰州附近了，车厢里隔壁铺的姑娘还在抓紧每分每秒学英语</p>\n<p><img src=\"https://aipaint.lihuanyu.com/%E7%81%AB%E8%BD%A6%E4%B8%8A%E5%AD%A6%E4%B9%A0%E7%9A%84%E5%A7%91%E5%A8%98.jpg\" alt=\"火车上学习的姑娘\"></p>\n<p>窗外的景色就和电视里对大西北的传统印象一样，黄土地、植物不多。</p>\n<p><img src=\"https://aipaint.lihuanyu.com/%E9%BB%84%E5%9C%9F%E5%9C%B0.jpeg\" alt=\"黄土地\"></p>\n<p>偶尔还能看到一些房子，已经被遗弃的样子，没有看到窑洞，估计还是房子住着更舒服一些吧。</p>\n<p><img src=\"https://aipaint.lihuanyu.com/%E9%BB%84%E5%9C%9F%E9%AB%98%E5%9D%A1%E4%B8%8A%E7%9A%84%E6%88%BF%E5%AD%90.jpeg\" alt=\"黄土高坡上的房子.jpeg\"></p>\n<p>相比于大漠风情，我还是更喜欢四川或者说南方那种植物茂密、郁郁葱葱的景象。一路过来有一种感觉就是，和小时候相比，农村里的人明显少了很多，几乎没有青壮年，都是老人家。</p>\n<p>随后火车在兰州东站停了大概1小时，再开十来分钟就到了兰州站。兰州站也是一个老站，没有现在那些高大上的新修的高铁站的现代感、科技感，却有一种历史的厚重感。</p>\n<p><img src=\"https://aipaint.lihuanyu.com/%E5%85%B0%E5%B7%9E%E4%B8%9C.jpeg\" alt=\"兰州东.jpeg\"></p>\n<p>到了兰州有新郎的朋友来接站，带着我吃了到兰州的第一顿早餐，兰州牛肉面。全国都能看到兰州拉面，但兰州并没有拉面，就像重庆没有鸡公煲。甚至重庆可能现在也有鸡公煲了，而兰州还找不到拉面馆，因为兰州都管这个叫兰州牛肉面。</p>\n<p><img src=\"https://aipaint.lihuanyu.com/%E5%85%B0%E5%B7%9E%E7%89%9B%E8%82%89%E9%9D%A2.jpeg\" alt=\"兰州牛肉面.jpeg\"></p>\n<p>大早上吃碗这个是真的舒服，很撑。住的酒店也挺不错，在黄河边上，可以直接看到黄河。</p>\n<p><img src=\"https://aipaint.lihuanyu.com/%E9%BB%84%E6%B2%B3.jpeg\" alt=\"黄河.jpeg\"></p>\n<p>黄河的水是真的黄，和电视里一样黄。以前去的景点也不少，但是还真少住在这种江河边可以直接看到江河的房子。所以对这个酒店是非常满意了，不过黄河确实小，还不如我家门口的绵远河宽。</p>\n<p>上午在酒店躺着休息了会儿，中午自己出去转悠，想起盗月社曾经在这边吃过一家烤串店，打车去吃了个午饭。</p>\n<div style=\"display: flex; justify-content: space-around\">\n    <img style=\"width: 50%\" src=\"https://aipaint.lihuanyu.com/烤串店.jpeg\" alt=\"烤串店\">\n    <img style=\"width: 50%\" src=\"https://aipaint.lihuanyu.com/烤牛肉.jpeg\" alt=\"烤牛肉\">\n</div>\n<p>看着还不错，吃着也好吃，价格不算便宜也不贵。单人吃下来48，20的肉20的筋，3块的饼5块的汽水。尤其是这饼，里面居然全是辣椒，把我一个四川人给整不会了，最后只吃了3/4。</p>\n<div style=\"display: flex; justify-content: space-around\">\n    <img style=\"width: 50%\" src=\"https://aipaint.lihuanyu.com/烤饼.jpeg\" alt=\"烤饼\">\n    <img style=\"width: 50%\" src=\"https://aipaint.lihuanyu.com/内含辣椒的烤饼.jpeg\" alt=\"内含辣椒的烤饼\">\n</div>\n<p>吃完继续溜达，兰州应该是一个三线城市，城建水平相比成都差了不少，和德阳感觉都差不多，堵车非常严重。兰州市中心的大广场还有当年的中国的味道</p>\n<p><img src=\"https://aipaint.lihuanyu.com/%E5%85%B0%E5%B7%9E%E7%9A%84%E5%A4%A7%E5%B9%BF%E5%9C%BA.jpeg\" alt=\"兰州的大广场.jpeg\"></p>\n<p>就在这个充斥红色氛围的广场旁边，就是新修的万象城大商城，看起来又非常的现代化。</p>\n<div style=\"display: flex; justify-content: space-around\">\n    <img style=\"width: 50%\" src=\"https://aipaint.lihuanyu.com/兰州万象城.jpeg\" alt=\"兰州万象城\">\n    <img style=\"width: 50%\" src=\"https://aipaint.lihuanyu.com/兰州万象城内部.jpeg\" alt=\"兰州万象城内部\">\n</div>\n<p>接下来是这个旅程中最满意的部分，兰州市博物馆。</p>\n<p><img src=\"https://aipaint.lihuanyu.com/%E5%85%B0%E5%B7%9E%E5%8D%9A%E7%89%A9%E9%A6%86.jpeg\" alt=\"兰州博物馆.jpeg\"></p>\n<p>里面藏品很多，也比较精美，应该是当年丝绸之路时代留下的文物。</p>\n<p><img src=\"https://aipaint.lihuanyu.com/%E5%85%B0%E5%B7%9E%E5%8D%9A%E7%89%A9%E9%A6%86%E8%97%8F%E5%93%81.jpeg\" alt=\"兰州博物馆藏品.jpeg\"></p>\n<p>这里面有个白衣寺塔的模型，很好看引起了我的注意</p>\n<p><img src=\"https://aipaint.lihuanyu.com/%E5%85%B0%E5%B7%9E%E5%8D%9A%E7%89%A9%E9%A6%86-%E7%99%BD%E8%A1%A3%E5%AF%BA%E5%A1%94.jpeg\" alt=\"兰州博物馆-白衣寺塔.jpeg\"></p>\n<p>主要是我越看这个模型越觉得眼熟，然后突然反应过来，这不是这个博物馆本身吗？一看介绍果然如此，这个白衣寺塔就是博物馆本身，寺庙建于明朝崇祯年，只有塔保留到了现在。博物馆是1991年迁入这里的。</p>\n<p>为什么我会反应过来这个就是博物馆本身呢？因为进来时候有注意到博物馆中央的大塔</p>\n<p><img src=\"https://aipaint.lihuanyu.com/%E5%85%B0%E5%B7%9E%E5%8D%9A%E7%89%A9%E9%A6%86-%E5%86%85%E9%83%A8%E7%99%BD%E8%A1%A3%E5%AF%BA%E5%A1%94.jpeg\" alt=\"兰州博物馆-内部白衣寺塔.jpeg\"></p>\n<p>里面还有左宗棠左公的雕像，还有贡院模型等，以及古代的兰州城沙盘</p>\n<p><img src=\"https://aipaint.lihuanyu.com/%E5%8F%A4%E4%BB%A3%E5%85%B0%E5%B7%9E%E6%B2%99%E7%9B%98.jpeg\" alt=\"古代兰州沙盘.jpeg\"></p>\n<p>逛完博物馆就回去休息了，晚上也只是在河边简单走了走。网上说兰州是白天阿富汗，夜间小香港，只能说晚上灯火是蛮好看的，小香港可能夸张了点。</p>\n<p>第二天参加婚礼，西北的婚礼和南方的婚礼又有一些不一样。同样是随份子钱，南方基本是收红包，红包上写名字，办完回去挨个拆红包记录礼金数量。北方就相对豪迈一些，参加的亲戚朋友都直接拿现金，不用红包，甚至可以电子支付，给了直接记录上。包了红包的也会现场拆开清点。给多给少都是心意，一些长辈混得不好的随200也很大方的给，混得好的给2000的也有也没有嘚瑟，所以和南方的方式比也没有什么好坏之分，只是地域不同习俗不同而已。</p>\n<p>室友本身就是一个很会搞事的，婚礼也很热闹，我本来以为他们会整点才艺展示，最后虽然是有表演，但并不是新郎伴郎表演的，而是专门的舞蹈队。一个比较有意思的是，大屏幕上青青草原，按我的思路应该会是一些比较柔和的音乐，实际确实像disco舞厅蹦迪一样的动次打次。</p>\n<p>给我的感觉是，围城无处不在。中原、江南等人口稠密区，本身就很热闹很繁华，所以这里的人们总想要一些隐居、清静的感觉，喜欢一种僻静的感觉。而西北大漠，虽然婚礼、聚会上很热闹，但我想，平时的他们可能没那么多人，所以才会更喜欢热闹吧。</p>\n<p>晚上再在市中心的步行街里的一家饭店里吃上一顿晚饭，就准备赶火车回家了。</p>\n<p>兰州的饭还有一个有意思的地方，上来会先给你端一杯茶，这个茶是按人头的，你要换位置请自己端好自己的茶。这个茶叫三炮台，有什么功能不清楚，但在这边吃了三顿比较正式的饭，每顿都会先给你一杯这个茶。</p>\n<p><img src=\"https://aipaint.lihuanyu.com/%E5%85%B0%E5%B7%9E-%E4%B8%89%E7%82%AE%E5%8F%B0%E8%8C%B6.jpeg\" alt=\"兰州-三炮台茶.jpeg\"></p>\n<p>最后走的时候，步行街上人来人往，我感觉这个城市的人是快乐的，比北京快乐，比成都快乐，而且孩子很多，在大街上三五成群打打闹闹，感觉这才是小孩应有的状态，而这个时间点，北上广的孩子应该都在写作业、补习班里吧。</p>\n<p><img src=\"https://aipaint.lihuanyu.com/%E5%85%B0%E5%B7%9E%E5%B8%82%E4%B8%AD%E5%BF%83%E6%AD%A5%E8%A1%8C%E8%A1%97.jpeg\" alt=\"兰州市中心步行街.jpeg\"></p>\n<p>总之，来过兰州了。你好兰州，再见兰州 👋🏻 。</p>\n","date_published":"2023-09-29T00:00:00.000Z","tags":["随笔","游记"],"language":"zh"},{"id":"https://www.lihuanyu.com/en/posts/2023/react-key-is-not-just-a-list-optimization/","url":"https://www.lihuanyu.com/en/posts/2023/react-key-is-not-just-a-list-optimization/","title":"React key Is Not Just a List Optimization","summary":"A practical explanation of how React key participates in component identity, why changing it resets state, and when that is the right tool.","content_html":"<p>While debugging a Mini Program form renderer, I ran into a familiar problem: the component kept too much internal state, and after switching to a different data object, the safest short-term fix was sometimes to destroy and recreate the component.</p>\n<p>In a Mini Program, the direct approach is usually conditional rendering. Make the component disappear, then show it again in the next tick. For example, set <code>a:if</code> to <code>false</code>, then set it back to <code>true</code>. That forces an unmount and a new creation.</p>\n<p>This naturally leads to a React comparison: in React, changing <code>key</code> can recreate a component. Does that mean <code>key</code> is a general refresh button?</p>\n<p>Not exactly. <code>key</code> is most often seen in list rendering, but in React its role is not only “helping list diffing.” It participates in deciding whether something is still the same component.</p>\n<p><a href=\"/posts/2023/React%E7%9A%84key/\">Chinese version of this article</a></p>\n<h2>key Participates in Component Identity</h2>\n<p>React associates state with a position in the render tree. In most cases, if the same component type stays in the same position, React preserves its state.</p>\n<p>When a component has a <code>key</code>, that key becomes part of its identity:</p>\n<pre><code class=\"language-jsx\">&lt;Editor key={articleId} article={article} /&gt;\n</code></pre>\n<p>When <code>articleId</code> changes, React does not treat this as “the same <code>Editor</code> received new props.” It treats it as “the old <code>Editor</code> was removed, and a new <code>Editor</code> was created.”</p>\n<p>The result is:</p>\n<ul>\n<li><code>useState</code> inside the component initializes again.</li>\n<li>State inside the child tree is reset too.</li>\n<li>Effects run through cleanup and setup as an unmount and mount.</li>\n<li>DOM nodes may be recreated instead of reused.</li>\n</ul>\n<p>So changing <code>key</code> is not an ordinary re-render. It is closer to replacing one component instance with another.</p>\n<h2>Re-rendering and Recreating Are Different</h2>\n<p>When a React component re-renders because props or state changed, its identity remains the same. Internal state is preserved. Effects run according to dependency changes. The DOM is reused where possible.</p>\n<p>When <code>key</code> changes, something else happens: React sees a different identity. The old component goes through unmounting, and the new component starts from its initial state.</p>\n<p>That is why changing <code>key</code> can reset a component. It does not refresh the component in place. It tells React to stop preserving the old subtree.</p>\n<h2>When key Is a Good Tool</h2>\n<p>The best case is when internal state should naturally belong to a business identity.</p>\n<p>For example, after switching contacts in a chat window, the draft for the previous contact should not remain in the input:</p>\n<pre><code class=\"language-jsx\">&lt;Chat key={to.id} contact={to} /&gt;\n</code></pre>\n<p>Or after switching articles in an editor, local draft state, validation state, and cursor-related state may all need to start from the new article:</p>\n<pre><code class=\"language-jsx\">&lt;ArticleEditor key={article.id} article={article} /&gt;\n</code></pre>\n<p>In these cases, using <code>key</code> is natural because the object in the user’s mind has changed. The component state should change with it.</p>\n<p>Forms, editors, wizards, and preview components often fall into this category. If the internal state belongs to a stable business object, using that object’s id as the <code>key</code> aligns component identity with business identity.</p>\n<h2>When key Is the Wrong Tool</h2>\n<p>Do not use <code>key</code> as a universal way to hide state bugs.</p>\n<p>If a component breaks unless it is destroyed and recreated, the relationship between internal state and external data is probably unclear. Maybe props changed but internal caches did not synchronize. Maybe a form renderer mixed schema, initial values, user input, and reset behavior into one state model. Changing <code>key</code> can make the symptom disappear, but it may only bypass the real issue.</p>\n<p>Avoid code like this:</p>\n<pre><code class=\"language-jsx\">&lt;Form key={Date.now()} /&gt;\n</code></pre>\n<p>or:</p>\n<pre><code class=\"language-jsx\">&lt;Form key={Math.random()} /&gt;\n</code></pre>\n<p>This makes React think the component identity changed on every render. Inputs lose content, focus disappears, effects run repeatedly, and performance gets worse.</p>\n<p>List keys follow the same principle. Prefer stable business ids. Array indexes are not always wrong, but if a list can insert, delete, or reorder items, index keys can make state follow the wrong item.</p>\n<h2>Back to the Form Renderer</h2>\n<p>For the Mini Program form renderer problem, changing <code>key</code> in React is indeed a possible tool. But its meaning is not “refresh this component.” Its meaning is “this is a different component now.”</p>\n<p>In Mini Program or Vue contexts, <code>key</code> does not behave exactly the same as React. If the goal is to force recreation, conditional rendering is often more direct. The more durable fix is to clean up the data flow so the component responds correctly when external data changes.</p>\n<p>If a form renderer must rely on destruction and recreation to work, it probably holds too much uncontrolled state. The long-term design should make several things explicit:</p>\n<ul>\n<li>What identifies the schema.</li>\n<li>How initial values differ from current user input.</li>\n<li>Who owns validation state, dirty state, and submitting state.</li>\n<li>Whether external data changes should synchronize existing state or explicitly trigger a reset.</li>\n</ul>\n<p>After these questions are clear, <code>key</code> can still be used for reset behavior. But it becomes an expression of identity change rather than a patch over state design problems.</p>\n<h2>Summary</h2>\n<p>React <code>key</code> is not just a list optimization. It participates in component identity. The question it answers is: is this component at this position still the same component?</p>\n<p>If the answer is no, a stable business id as <code>key</code> can reset state cleanly. If the only reason to change <code>key</code> is that internal state has become hard to reason about, the data flow and state boundaries need another look.</p>\n<p>The React documentation has two related sections:</p>\n<ul>\n<li><a href=\"https://react.dev/learn/preserving-and-resetting-state\">Preserving and Resetting State</a></li>\n<li><a href=\"https://react.dev/reference/react/useState#resetting-state-with-a-key\">Resetting state with a key</a></li>\n</ul>\n","date_published":"2023-09-19T00:00:00.000Z","date_modified":"2026-05-05T00:00:00.000Z","tags":["Frontend","React"],"language":"en"},{"id":"https://www.lihuanyu.com/posts/2023/React%E7%9A%84key/","url":"https://www.lihuanyu.com/posts/2023/React%E7%9A%84key/","title":"React 里的 key 不只是列表优化","summary":"从一次组件重建问题出发，解释 React key 如何参与组件身份判断，以及什么时候适合用 key 重置状态。","content_html":"<p>一次排查小程序表单渲染器问题时，遇到一个很典型的场景：组件内部状态处理得不够干净，切换数据后偶尔需要销毁再创建。</p>\n<p>在小程序里，直接的做法通常是条件渲染：先让组件消失，下一轮再让它出现。比如先把 <code>a:if</code> 改成 <code>false</code>，再改回 <code>true</code>。这样可以触发卸载和重新创建。</p>\n<p>这个问题很容易让人联想到 React：如果想让一个组件重新创建，改一下 <code>key</code> 不就行了吗？</p>\n<p>这个说法没错，但容易让人误解 <code>key</code> 的作用。<code>key</code> 最常见的出现场景确实是列表渲染，但它在 React 里的意义不只是“辅助列表 diff”，而是参与判断“这是不是同一个组件”。</p>\n<p><a href=\"/en/posts/2023/react-key-is-not-just-a-list-optimization/\">English version: React key Is Not Just a List Optimization</a></p>\n<h2>key 参与的是组件身份判断</h2>\n<p>React 会把状态绑定到渲染树里的某个位置。多数情况下，只要同一个组件类型还在同一个位置，React 就会保留它的状态。</p>\n<p>如果给组件加上 <code>key</code>，这个 <code>key</code> 就会成为组件身份的一部分：</p>\n<pre><code class=\"language-jsx\">&lt;Editor key={articleId} article={article} /&gt;\n</code></pre>\n<p>当 <code>articleId</code> 改变时，React 不会把它理解成“同一个 <code>Editor</code> 组件收到了一份新 props”，而是理解成“旧的 <code>Editor</code> 被移除，新的 <code>Editor</code> 被创建”。</p>\n<p>结果就是：</p>\n<ul>\n<li>组件内部的 <code>useState</code> 会重新初始化；</li>\n<li>子组件树里的状态也会一起重置；</li>\n<li>effect 的清理和重新执行也会按卸载、挂载流程走；</li>\n<li>DOM 节点可能会被重新创建，而不是继续复用。</li>\n</ul>\n<p>所以，改 <code>key</code> 不是普通意义上的“重新渲染”，而是更接近“换了一个组件实例”。</p>\n<h2>重新渲染和重新创建不是一回事</h2>\n<p>React 组件因为 props 或 state 变化而重新渲染时，组件实例的身份没有变。组件内部的 state 会被保留，effect 会按照依赖变化决定是否重新执行，DOM 也会尽量复用。</p>\n<p><code>key</code> 改变时发生的是另一件事：React 认为组件身份变了。旧组件会走卸载流程，新组件会从初始状态开始。</p>\n<p>这也是为什么改 <code>key</code> 常常可以“重置”组件。它不是让组件刷新一下，而是让 React 放弃原来的那棵子树。</p>\n<h2>什么时候适合用</h2>\n<p>最适合的场景是：组件内部状态本来就应该跟某个业务身份绑定。</p>\n<p>比如聊天窗口切换联系人后，不希望把上一个人的输入草稿保留下来：</p>\n<pre><code class=\"language-jsx\">&lt;Chat key={to.id} contact={to} /&gt;\n</code></pre>\n<p>或者编辑器切换文章后，希望本地草稿、校验状态、光标位置都从新文章开始：</p>\n<pre><code class=\"language-jsx\">&lt;ArticleEditor key={article.id} article={article} /&gt;\n</code></pre>\n<p>这种时候用 <code>key</code> 很自然，因为用户理解里的“对象”已经变了，组件状态也应该跟着换一份。</p>\n<p>表单、编辑器、向导页、预览器这类组件尤其常见。只要内部状态属于某个稳定业务对象，用这个对象的 id 作为 <code>key</code>，就是在把组件身份和业务身份对齐。</p>\n<h2>什么时候不该用</h2>\n<p>不应该把 <code>key</code> 当成修复状态 bug 的万能开关。</p>\n<p>如果组件只要不销毁重建就会出错，通常说明内部状态和外部数据的关系没有理顺。比如 props 变了，但组件内部缓存没有同步；或者表单渲染器把 schema、默认值、用户输入混在了一起。这个时候改 <code>key</code> 能让问题消失，但也可能只是绕开了真正的问题。</p>\n<p>尤其不要写这种代码：</p>\n<pre><code class=\"language-jsx\">&lt;Form key={Date.now()} /&gt;\n</code></pre>\n<p>或者：</p>\n<pre><code class=\"language-jsx\">&lt;Form key={Math.random()} /&gt;\n</code></pre>\n<p>这样每次渲染都会让 React 认为组件身份变了。输入框会丢内容，焦点会丢失，effect 会反复执行，性能也会变差。</p>\n<p>列表里的 <code>key</code> 也一样，应该优先使用稳定的业务 id。用数组下标做 <code>key</code> 不是绝对不行，但如果列表会插入、删除、排序，就很容易让状态跟错项目。</p>\n<h2>表单渲染器的问题怎么处理</h2>\n<p>回到开头的小程序表单渲染器问题，React 里改 <code>key</code> 确实是一种可用手段，但它背后的语义不是“刷新一下组件”，而是“告诉 React 这是另一个组件”。</p>\n<p>在小程序或 Vue 的上下文里，<code>key</code> 的行为和 React 不完全一样。要强制重建组件，条件渲染是更直接的办法；更稳妥的做法则是把组件的数据流整理清楚，让它能正确响应外部数据变化。</p>\n<p>如果一个表单渲染器必须靠销毁重建才能工作，大概率是组件内部藏了太多不受控状态。短期可以用重建救急，长期还是应该把几件事拆清楚：</p>\n<ul>\n<li>schema 的身份是什么；</li>\n<li>初始值和用户当前输入如何区分；</li>\n<li>校验状态、脏状态和提交状态归谁管理；</li>\n<li>外部数据变化时，是同步已有状态，还是明确执行一次 reset。</li>\n</ul>\n<p>这些问题想清楚后，<code>key</code> 仍然可以作为重置手段，但它不再是掩盖状态设计问题的补丁，而是一个表达组件身份变化的工具。</p>\n<h2>小结</h2>\n<p>React 里的 <code>key</code> 不只是列表优化。它参与组件身份判断，回答的问题是：当前位置上的组件，还是不是同一个组件？</p>\n<p>如果答案是否定的，用稳定的业务 id 作为 <code>key</code> 可以自然地重置状态。如果只是因为组件内部状态处理混乱而想强制销毁重建，就应该先回头看数据流和状态边界。</p>\n<p>React 官方文档里有两处相关说明：</p>\n<ul>\n<li><a href=\"https://react.dev/learn/preserving-and-resetting-state\">Preserving and Resetting State</a></li>\n<li><a href=\"https://react.dev/reference/react/useState#resetting-state-with-a-key\">Resetting state with a key</a></li>\n</ul>\n","date_published":"2023-09-19T00:00:00.000Z","date_modified":"2026-05-05T00:00:00.000Z","tags":["前端","React"],"language":"zh"},{"id":"https://www.lihuanyu.com/posts/2023/Mac%E7%A3%81%E7%9B%98%E6%B8%85%E7%90%86%E5%B7%A5%E5%85%B7%E6%8E%A8%E8%8D%90/","url":"https://www.lihuanyu.com/posts/2023/Mac%E7%A3%81%E7%9B%98%E6%B8%85%E7%90%86%E5%B7%A5%E5%85%B7%E6%8E%A8%E8%8D%90/","title":"Mac 磁盘清理工具推荐：先看清楚，再决定删什么","summary":"推荐 OmniDiskSweeper 作为 Mac 磁盘空间分析工具，并补充 Apple 自带存储管理、DaisyDisk、GrandPerspective 等替代方案，以及清理缓存、系统数据和大文件时的风险边界。","content_html":"<p>Mac 电脑有一个很苹果的特点：机器很好，内存和硬盘也很好，只是买的时候每往上加一点容量，价格都让人清醒。</p>\n<p>刚买的时候觉得 512GB 够用。用几年以后，Xcode、Docker、微信、照片、视频、缓存、日志、各种 SDK 一起长大，硬盘就开始报警。系统设置里一看，几十上百 GB 都变成了“系统数据”或者某个不太说人话的分类。</p>\n<p>这时候最需要的不是立刻清理，而是先看清楚。</p>\n<p>磁盘清理工具分两种。一种是“自动一键清理”，听起来省事，但风险也在这里；另一种是“把空间占用列出来，用户自己决定删什么”。我更信任后者。</p>\n<h2>先看 macOS 自带工具</h2>\n<p>macOS 自己就有存储管理入口。</p>\n<p>在 macOS Ventura 13 及以后，可以从“系统设置 - 通用 - 储存空间”查看。Apple 官方文档也列了几类常见处理方式：查看可用空间、优化存储、移动或删除文件、清理下载、删除旧备份、卸载不用的应用、清空废纸篓等。</p>\n<p>官方页面在这里：<a href=\"https://support.apple.com/en-us/102624\">Free up storage space on Mac</a></p>\n<p>系统自带工具的优点是安全，缺点是粗。</p>\n<p>它能告诉你大概是应用、文档、照片、信息、音乐、废纸篓占了空间，但遇到“系统数据”这类大桶时，经常只能看个热闹。系统数据并不全是系统文件，它可能包含缓存、日志、iOS 备份、Time Machine 本地快照、应用支持文件、开发工具产物，以及各种不容易归类的东西。</p>\n<p>这时就需要磁盘分析工具了。</p>\n<h2>OmniDiskSweeper：朴素，但够用</h2>\n<p>我一直比较喜欢 <a href=\"https://www.omnigroup.com/blog/omnidisksweeper-1.10\">OmniDiskSweeper</a>。</p>\n<p>它属于 The Omni Group 的小工具。官方介绍很直接：帮你找到 Mac 上可以删除的文件，从而释放磁盘空间。Omni 在 2018 年还更新过 1.10 版本，说明它不是完全被遗忘的古董，只是确实不怎么折腾。</p>\n<p>OmniDiskSweeper 的界面很朴素。</p>\n<p>它会按目录大小排序，一层层展开，让你看到空间到底被谁吃掉了。它不假装自己很聪明，也不弹一堆吓人的提示，更不会告诉你“一键优化 30GB”。它只是把文件大小摆出来，删不删由你决定。</p>\n<p>这正是我喜欢它的原因。</p>\n<p>磁盘清理最怕工具太积极。电脑里很多东西看起来像垃圾，实际删掉就会出事。缓存可以重建，但重建也要时间；日志通常能删，但不代表所有日志都无用；应用支持目录里可能有缓存，也可能有项目数据。一个工具如果一上来就替用户做判断，看起来贴心，实际像拿着扫帚进仓库，见东西就扫。</p>\n<p>OmniDiskSweeper 更像手电筒。</p>\n<p>它负责照亮角落，不替你动手。</p>\n<p><img src=\"https://aipaint.lihuanyu.com/Omnisweeper%E7%95%8C%E9%9D%A2.png\" alt=\"Omnisweeper界面\"></p>\n<h2>还有几个替代工具</h2>\n<p>如果只想免费、简单、看目录大小，OmniDiskSweeper 够用。</p>\n<p>如果希望图形化更强，可以看 <a href=\"https://grandperspectiv.sourceforge.net/\">GrandPerspective</a>。它用矩形树图展示磁盘占用，适合一眼看出哪个文件或目录特别大。官方页面也说明，它就是一个用来图形化展示 macOS 文件系统磁盘占用的小工具。</p>\n<p>如果愿意付费，并且想要更好的界面、速度和隐藏空间提示，可以看 <a href=\"https://daisydiskapp.com/\">DaisyDisk</a>。它支持扫描本地、外接、网络和部分云盘来源，也会处理隐藏空间、可清理空间等问题。官方文档里也强调，它不会自动清理，最终还是用户自己决定删除哪些文件。</p>\n<p>这几个工具的取向大概是：</p>\n<ol>\n<li>OmniDiskSweeper：列表清晰，免费，适合直接找大目录。</li>\n<li>GrandPerspective：免费，图形化强，适合快速定位异常大文件。</li>\n<li>DaisyDisk：付费，体验好，适合经常清理或希望看隐藏空间的人。</li>\n<li>系统自带存储管理：最安全，适合先做基础检查。</li>\n</ol>\n<p>我不太推荐那种主打“一键清理、系统加速、深度优化”的工具。</p>\n<p>不是说它们一定有问题，而是这类软件的商业模式常常鼓励它们制造焦虑。缓存不等于垃圾，内存占用不等于浪费，系统数据不等于毒瘤。Mac 不需要每天被体检，电脑也不是天天要排毒的人体。</p>\n<p>真正值得删的，往往不是工具吓出来的东西，而是自己确实不再需要的东西。</p>\n<h2>哪些东西通常值得看</h2>\n<p>对开发者来说，Mac 磁盘爆掉，常见来源有这些：</p>\n<ol>\n<li><code>~/Downloads</code>：下载过的安装包、压缩包、视频、临时文件。</li>\n<li>Xcode DerivedData：旧项目编译缓存可能很大。</li>\n<li>iOS Simulator：不用的模拟器和运行时。</li>\n<li>Docker images、containers、volumes：旧镜像和卷很容易堆起来。</li>\n<li><code>node_modules</code>：多个项目累积起来很夸张。</li>\n<li>pnpm、npm、yarn 缓存：通常可以清，但要知道清完会重新下载。</li>\n<li>微信、飞书、钉钉等 IM 缓存：聊天文件和图片会一直长。</li>\n<li>旧 iPhone/iPad 备份：一份备份几十 GB 很常见。</li>\n<li>视频剪辑、录屏、设计素材：单个文件可能巨大。</li>\n<li>废弃 SDK、旧数据库 dump、测试数据。</li>\n</ol>\n<p>这些地方的共同点是：它们大多属于用户或开发环境，不是系统核心。</p>\n<p>清理时可以遵循一个原则：先删自己认识的东西，再删工具建议的东西。</p>\n<p>一个文件如果不知道是什么，不要因为它大就删。大文件不一定是垃圾，小文件也不一定安全。磁盘空间不够很烦，但系统坏掉更烦。</p>\n<h2>哪些地方不要乱碰</h2>\n<p>有些目录最好不要靠感觉删。</p>\n<p>比如 <code>/System</code>、<code>/Library</code>、<code>/usr</code>、<code>/private</code>、<code>/var</code>、<code>/bin</code>、<code>/sbin</code> 这类路径。里面当然也可能有大文件，但普通用户不应该把它们当清理对象。很多应用依赖、系统服务、权限数据、临时文件和运行状态都在这些地方。</p>\n<p>用户目录下也不是所有东西都能随便删。</p>\n<p><code>~/Library/Application Support</code> 里有很多应用数据。某个应用目录特别大时，要先判断里面是缓存、下载内容，还是用户数据。比如一个编辑器、数据库工具、虚拟机、笔记软件、设计软件，都可能把重要数据放在这里。</p>\n<p>还有 APFS 的本地快照和 purgeable space。</p>\n<p>有时 Finder、系统设置、第三方工具显示的空间不一致，不一定是工具错，也不一定是系统坏。APFS、Time Machine、本地快照、云盘占位文件都会让“空间到底被谁用了”变得没那么直观。遇到这种情况，先重启、清空废纸篓、检查 Time Machine，再考虑更进一步处理。</p>\n<p>不要因为看到“系统数据 200GB”就冲动。</p>\n<p>这个桶很讨厌，但它不是一个文件夹。它更像一个筐，macOS 把不好归类的东西都往里放。要处理它，还是要回到具体目录和具体文件。</p>\n<h2>我的清理顺序</h2>\n<p>如果现在 Mac 提示空间不足，我通常会这样做：</p>\n<ol>\n<li>先看系统设置里的储存空间，确认大类。</li>\n<li>清空废纸篓和下载目录里的明确废弃文件。</li>\n<li>用 OmniDiskSweeper 或 DaisyDisk 扫用户目录。</li>\n<li>先处理自己认识的大文件和大目录。</li>\n<li>再看开发工具缓存、Docker、模拟器、旧备份。</li>\n<li>不认识的系统目录先搜索确认，不急着删。</li>\n<li>清完后重启一次，再看空间是否释放。</li>\n</ol>\n<p>如果只是想在命令行里粗略看目录大小，也可以用：</p>\n<pre><code class=\"language-bash\">du -sh * | sort -h\n</code></pre>\n<p>或者安装 <code>ncdu</code> 这类命令行工具。只是命令行更锋利，删东西前更要看清楚路径。</p>\n<h2>最后还是那句话</h2>\n<p>磁盘清理不是清洁大扫除，更像盘账。</p>\n<p>账要先看明白，再决定哪笔该处理。一个好的磁盘工具，不应该替用户表演勤快，而应该把事实摆清楚：谁占了空间，占了多少，在哪里。</p>\n<p>OmniDiskSweeper 的好处就在这里。它不华丽，也不热闹，但干净。</p>\n<p>Mac 空间不够时，最该警惕的不是大文件，而是自己不知道自己在删什么。</p>\n<p>先看清楚，再决定删什么。</p>\n","date_published":"2023-09-16T00:00:00.000Z","date_modified":"2026-05-16T00:00:00.000Z","tags":["mac","工具","磁盘清理","mac磁盘清理"],"language":"zh"},{"id":"https://www.lihuanyu.com/en/posts/2023/sql-injection-investigation-nestjs/","url":"https://www.lihuanyu.com/en/posts/2023/sql-injection-investigation-nestjs/","title":"A SQL Injection Incident Review: NestJS Validation, Logs, and Server-Side Security","summary":"A practical review of a SQL injection issue found during a Mini Program security test, covering NestJS validation, ORM query safety, PM2 logs, database constraints, and defense in depth.","content_html":"<p>I once built a Mini Program with a NestJS backend and MySQL database. During submission review, the platform offered an API security test that simulated common attack requests.</p>\n<p>The test did not cause real damage, but it inserted dozens of unexpected blank records into the database. The issue was small, but it was a useful reminder: this was not just one missing <code>parseInt</code>. Several layers of the server-side safety net were incomplete.</p>\n<p><a href=\"/posts/2023/%E8%AE%B0%E5%BD%95%E4%B8%80%E6%AC%A1SQL%E6%B3%A8%E5%85%A5%E4%B8%8E%E9%97%AE%E9%A2%98%E6%8E%92%E6%9F%A5/\">Chinese version of this article</a></p>\n<h2>How the Problem Was Found</h2>\n<p>During Mini Program review, the platform showed an interface security test option:</p>\n<p><img src=\"https://aipaint.lihuanyu.com/66dca5b4a7937737a9bb6d68a1e29ff.png\" alt=\"Mini Program interface security test\"></p>\n<p>My first reaction was that the risk should be low. The server was hand-written, not an old open-source system with known historical vulnerabilities. The database only accepted local connections. The API had input and output validation. It felt safe enough.</p>\n<p>After the test started, the server logs showed many requests, but no obvious errors. The real problem surfaced later in the admin page: the history list contained dozens of blank records. The database confirmed that those rows were not created by the normal business flow.</p>\n<p><img src=\"https://aipaint.lihuanyu.com/%E6%95%B0%E6%8D%AE%E5%BA%93%E8%A2%AB%E6%B3%A8%E5%85%A5%E5%90%8E%E7%9A%84%E5%BC%82%E5%B8%B8%E6%95%B0%E6%8D%AE.png\" alt=\"abnormal database rows\"></p>\n<p>For this kind of incident, two questions matter immediately:</p>\n<ul>\n<li>Was this only junk data being written?</li>\n<li>Was there any unauthorized read, bulk delete, data leak, or privilege escalation?</li>\n</ul>\n<p>In this case, I only found abnormal writes and did not see evidence of sensitive data leakage. But from a security perspective, once an attack payload can affect SQL semantics, it should not be treated as a harmless data cleanup issue.</p>\n<h2>Locating the Entry Point with PM2 Logs</h2>\n<p>The service was deployed with PM2. Real-time logs are available through:</p>\n<pre><code class=\"language-bash\">pm2 logs\n</code></pre>\n<p>For historical investigation, PM2’s log files on disk are more useful. By default, PM2 saves logs under:</p>\n<pre><code class=\"language-text\">$HOME/.pm2/logs\n</code></pre>\n<p>After downloading the relevant <code>out</code> and <code>error</code> logs, <code>rg</code> is enough for a first pass:</p>\n<pre><code class=\"language-bash\">rg -n -i &quot;union select|sleep\\\\(|or 1=1|--|/\\\\*&quot; app-out.log\n</code></pre>\n<p>The logs showed a request similar to this:</p>\n<pre><code class=\"language-text\">User 1676 requested history image list page 1&quot; union select 1,2--\n</code></pre>\n<p>The original log also contained terminal color control characters:</p>\n<p><img src=\"https://aipaint.lihuanyu.com/sql%E6%B3%A8%E5%85%A5%E6%97%A5%E5%BF%97%E5%88%86%E6%9E%90.png\" alt=\"SQL injection log analysis\"></p>\n<p>Those were ANSI escape codes, not an encoding problem. When needed, the log can be cleaned before reading:</p>\n<pre><code class=\"language-bash\">perl -pe 's/\\e\\[[0-9;]*[mK]//g' app-out.log &gt; app-out.clean.log\n</code></pre>\n<p>The entry point was the pagination parameter of the history list API. The endpoint expected a page number, but the page number position received a SQL injection payload.</p>\n<h2>Root Cause: Treating a Query Parameter as a Trusted Number</h2>\n<p>A pagination endpoint usually looks like this:</p>\n<pre><code class=\"language-http\">GET /histories?page=1&amp;pageSize=20\n</code></pre>\n<p>The problem was that the server used <code>page</code> as a number, while HTTP query parameters always arrive as strings. Without runtime conversion and validation, <code>page</code> can be any string.</p>\n<p>Two dangerous patterns are common.</p>\n<p>The first is raw SQL string interpolation:</p>\n<pre><code class=\"language-ts\">const sql = `\n  SELECT * FROM histories\n  WHERE user_id = ${userId}\n  ORDER BY created_at DESC\n  LIMIT ${(page - 1) * pageSize}, ${pageSize}\n`;\n</code></pre>\n<p>The second uses an ORM, but still interpolates strings in parts of the query:</p>\n<pre><code class=\"language-ts\">queryBuilder\n  .where(`history.user_id = ${userId}`)\n  .take(pageSize)\n  .skip((page - 1) * pageSize);\n</code></pre>\n<p>Both patterns put untrusted input into SQL structure. If the input is not constrained to a real number, it may change the meaning of the query.</p>\n<p>The fix should not be “filter this specific payload”. SQL injection is not about a few dangerous strings. It happens when user input is executed as part of SQL code.</p>\n<p>OWASP’s primary defense for SQL injection is parameterized queries: define SQL structure first, then bind user input as data so the database can distinguish code from values. Input validation is also important, but it is not a replacement for parameterized queries.</p>\n<h2>Validating Parameters in NestJS</h2>\n<p>NestJS provides Pipes and <code>ValidationPipe</code>, which are good places to enforce request boundaries.</p>\n<p>For simple pagination, built-in pipes are enough:</p>\n<pre><code class=\"language-ts\">import { DefaultValuePipe, ParseIntPipe, Query } from '@nestjs/common';\n\n@Get('histories')\nasync listHistories(\n  @Query('page', new DefaultValuePipe(1), ParseIntPipe) page: number,\n  @Query('pageSize', new DefaultValuePipe(20), ParseIntPipe) pageSize: number,\n) {\n  const safePage = Math.max(page, 1);\n  const safePageSize = Math.min(Math.max(pageSize, 1), 50);\n\n  return this.historyService.list({\n    page: safePage,\n    pageSize: safePageSize,\n  });\n}\n</code></pre>\n<p>For larger query objects, a DTO is cleaner:</p>\n<pre><code class=\"language-ts\">import { Type } from 'class-transformer';\nimport { IsInt, Max, Min } from 'class-validator';\n\nexport class ListHistoryQueryDto {\n  @Type(() =&gt; Number)\n  @IsInt()\n  @Min(1)\n  page = 1;\n\n  @Type(() =&gt; Number)\n  @IsInt()\n  @Min(1)\n  @Max(50)\n  pageSize = 20;\n}\n</code></pre>\n<p>Enable <code>ValidationPipe</code> globally:</p>\n<pre><code class=\"language-ts\">app.useGlobalPipes(\n  new ValidationPipe({\n    transform: true,\n    whitelist: true,\n    forbidNonWhitelisted: true,\n  }),\n);\n</code></pre>\n<p>Several details matter:</p>\n<ul>\n<li><code>transform: true</code> lets DTOs convert query strings to numbers.</li>\n<li><code>whitelist: true</code> strips fields not declared in the DTO.</li>\n<li><code>forbidNonWhitelisted: true</code> rejects extra fields instead of silently dropping them.</li>\n<li><code>@Type(() =&gt; Number)</code>, <code>@IsInt()</code>, <code>@Min()</code>, and <code>@Max()</code> should work together. A TypeScript <code>number</code> type alone is not runtime validation.</li>\n</ul>\n<p>TypeScript types disappear at runtime. HTTP requests still arrive as strings. Server boundaries need runtime validation.</p>\n<h2>ORM Queries Still Need Safe APIs</h2>\n<p>Using an ORM does not automatically eliminate SQL injection. It depends on whether the ORM’s parameter binding features are used correctly.</p>\n<p>A repository API is usually safer:</p>\n<pre><code class=\"language-ts\">return this.historyRepository.find({\n  where: {\n    userId,\n  },\n  order: {\n    createdAt: 'DESC',\n  },\n  skip: (page - 1) * pageSize,\n  take: pageSize,\n});\n</code></pre>\n<p>If QueryBuilder is necessary, bind parameters:</p>\n<pre><code class=\"language-ts\">return this.historyRepository\n  .createQueryBuilder('history')\n  .where('history.user_id = :userId', { userId })\n  .orderBy('history.created_at', 'DESC')\n  .skip((page - 1) * pageSize)\n  .take(pageSize)\n  .getMany();\n</code></pre>\n<p>Avoid this:</p>\n<pre><code class=\"language-ts\">.where(`history.user_id = ${userId}`)\n</code></pre>\n<p>Dynamic sorting is another common trap. Table names, column names, and sort directions are SQL structure and often cannot be handled with normal value binding. Use an allow-list:</p>\n<pre><code class=\"language-ts\">const sortFields = {\n  createdAt: 'history.created_at',\n  id: 'history.id',\n} as const;\n\nconst sortDirections = {\n  asc: 'ASC',\n  desc: 'DESC',\n} as const;\n\nconst sortField = sortFields[query.sortBy] ?? sortFields.createdAt;\nconst sortDirection = sortDirections[query.order] ?? sortDirections.desc;\n\nqueryBuilder.orderBy(sortField, sortDirection);\n</code></pre>\n<p>The user input is not inserted into SQL. It is mapped to server-defined safe choices.</p>\n<h2>Database Constraints Are the Last Reminder</h2>\n<p>The admin page showed blank records, which means the database layer was also missing useful constraints.</p>\n<p>Fields that are required by business logic should also be expressed in the database:</p>\n<ul>\n<li><code>NOT NULL</code></li>\n<li>Reasonable <code>VARCHAR</code> length</li>\n<li>Enum or status constraints</li>\n<li>Foreign keys or logical foreign keys</li>\n<li><code>created_at</code> and <code>updated_at</code> defaults</li>\n<li>Necessary unique indexes</li>\n</ul>\n<p>Database constraints do not replace server-side validation, and they do not prevent SQL injection by themselves. But when the server misses a boundary, constraints can turn “silent bad data” into a failed write and an alert.</p>\n<p>For a history table, if user ID, image URL, status, and creation time are required, blank rows should not be insertable.</p>\n<h2>Least Privilege Still Matters</h2>\n<p>The damage caused by SQL injection depends heavily on the database account’s permissions.</p>\n<p>Small projects often let the application use a powerful database account, sometimes one that can create tables, drop tables, or change schema. It is convenient, but it expands the blast radius.</p>\n<p>A safer approach is:</p>\n<ul>\n<li>The runtime application account only has the required <code>SELECT</code>, <code>INSERT</code>, <code>UPDATE</code>, and <code>DELETE</code> permissions.</li>\n<li>Migration credentials and runtime credentials are separated.</li>\n<li>The application account does not get <code>DROP</code>, <code>ALTER</code>, or full database administration privileges.</li>\n<li>Different services or databases use different accounts where practical.</li>\n</ul>\n<p>OWASP also lists least privilege as defense in depth for SQL injection. It does not prevent the bug, but it reduces the impact after a successful exploit.</p>\n<h2>Logs Should Help Investigation</h2>\n<p>The issue was found because the logs contained the suspicious request. But the original log format still had several weaknesses:</p>\n<ul>\n<li>No consistent request ID.</li>\n<li>User, endpoint, and parameters were not structured.</li>\n<li>Terminal color codes were mixed into persisted logs.</li>\n<li>Suspicious requests did not trigger separate alerts.</li>\n</ul>\n<p>Server logs should be able to answer:</p>\n<ul>\n<li>Which user or anonymous identifier made the request?</li>\n<li>What were the method and route?</li>\n<li>What key query/body parameters were sent, with sensitive fields redacted?</li>\n<li>What was the response status and latency?</li>\n<li>How does an exception stack trace connect to request context?</li>\n</ul>\n<p>Structured logs are better than colored text for production investigation. Even without a full logging platform, JSON logs are easier to process later with <code>rg</code>, <code>jq</code>, Loki, or ELK.</p>\n<h2>Platform Testing Is Not Enough</h2>\n<p>The platform’s simulated attack was valuable because it exposed the issue. But external security testing should be treated as a signal, not as the main security system.</p>\n<p>At least a few tests should be added:</p>\n<pre><code class=\"language-ts\">it('rejects non-numeric page query', async () =&gt; {\n  await request(app.getHttpServer())\n    .get('/histories?page=1%22%20union%20select%201,2--')\n    .expect(400);\n});\n\nit('limits pageSize', async () =&gt; {\n  await request(app.getHttpServer())\n    .get('/histories?page=1&amp;pageSize=10000')\n    .expect(400);\n});\n</code></pre>\n<p>Service-level tests can verify that pagination values pass through DTOs or pipes before reaching query logic. An integration test can also confirm that invalid requests do not insert business records.</p>\n<p>Security testing does not need to be complex at the start. Turning known failures into regression tests is the highest-return step.</p>\n<h2>A Server-Side Security Review Checklist</h2>\n<p>After this incident, I would check similar projects with this list:</p>\n<ol>\n<li>Do all <code>params</code>, <code>query</code>, and <code>body</code> inputs have runtime validation?</li>\n<li>Are pagination parameters integers with minimum and maximum bounds?</li>\n<li>Are dynamic sort fields mapped through allow-lists?</li>\n<li>Do ORM queries use parameter binding, and are any raw SQL strings still interpolating user input?</li>\n<li>Do database fields have required <code>NOT NULL</code>, length, index, and status constraints?</li>\n<li>Does the production database user follow least privilege?</li>\n<li>Do error responses avoid leaking SQL, table names, stack traces, and internal paths?</li>\n<li>Can logs correlate request ID, user, endpoint, parameters, status, and exception?</li>\n<li>Are known attack payloads covered by e2e tests?</li>\n<li>Are there alerts for abnormal writes, error spikes, and unusual 400/500 patterns?</li>\n</ol>\n<p>SQL injection is rarely an isolated line-level bug. It often points to unclear input boundaries, unsafe query construction, weak database constraints, and poor observability at the same time.</p>\n<h2>Summary</h2>\n<p>The direct fix was to convert and validate pagination parameters. But the real lesson is broader.</p>\n<p>Server-side security needs layers:</p>\n<ul>\n<li>Controllers use pipes and DTOs for runtime validation.</li>\n<li>Query code uses parameterized queries and safe ORM APIs.</li>\n<li>Dynamic SQL structure uses allow-lists.</li>\n<li>The database uses constraints and least privilege to reduce impact.</li>\n<li>Logs and tests make the issue easier to detect and reproduce.</li>\n</ul>\n<p>The platform test only brought the problem into view. The system becomes safer only when the incident is converted into code constraints, database constraints, tests, and investigation workflow.</p>\n<h2>Further Reading</h2>\n<ul>\n<li><a href=\"https://cheatsheetseries.owasp.org/cheatsheets/SQL_Injection_Prevention_Cheat_Sheet.html\">OWASP: SQL Injection Prevention Cheat Sheet</a></li>\n<li><a href=\"https://cheatsheetseries.owasp.org/cheatsheets/Input_Validation_Cheat_Sheet.html\">OWASP: Input Validation Cheat Sheet</a></li>\n<li><a href=\"https://docs.nestjs.com/techniques/validation\">NestJS: Validation</a></li>\n<li><a href=\"https://docs.nestjs.com/pipes\">NestJS: Pipes</a></li>\n<li><a href=\"https://typeorm.io/docs/query-builder/select-query-builder\">TypeORM: Select using Query Builder</a></li>\n<li><a href=\"https://pm2.io/docs/runtime/guide/log-management/\">PM2: Log Management</a></li>\n</ul>\n","date_published":"2023-08-20T00:00:00.000Z","date_modified":"2026-05-04T00:00:00.000Z","tags":["SQL Injection","NestJS","MySQL","Server-Side Security","PM2 Logs"],"language":"en"},{"id":"https://www.lihuanyu.com/posts/2023/%E8%AE%B0%E5%BD%95%E4%B8%80%E6%AC%A1SQL%E6%B3%A8%E5%85%A5%E4%B8%8E%E9%97%AE%E9%A2%98%E6%8E%92%E6%9F%A5/","url":"https://www.lihuanyu.com/posts/2023/%E8%AE%B0%E5%BD%95%E4%B8%80%E6%AC%A1SQL%E6%B3%A8%E5%85%A5%E4%B8%8E%E9%97%AE%E9%A2%98%E6%8E%92%E6%9F%A5/","title":"一次 SQL 注入排查复盘：NestJS、日志与服务端安全","summary":"从一次小程序安全测试触发的 SQL 注入排查出发，复盘 NestJS 参数校验、ORM 查询写法、日志定位、数据库约束和服务端安全防线。","content_html":"<p>之前做过一个小程序，服务端用 NestJS，数据库是 MySQL。小程序提审时，微信平台提供了一次接口安全测试，模拟用户请求去探测常见漏洞。</p>\n<p>测试本身没有造成真实破坏，但数据库里被写入了几十条不符合预期的空白记录。这个问题影响范围不大，却很适合复盘：它暴露的不是某一行 <code>parseInt</code> 忘了写，而是服务端安全里几层防线都不够完整。</p>\n<p><a href=\"/en/posts/2023/sql-injection-investigation-nestjs/\">English version: A SQL Injection Incident Review: NestJS Validation, Logs, and Server-Side Security</a></p>\n<h2>问题是怎么发现的</h2>\n<p>小程序提审时，微信后台提示可以进行接口安全测试：</p>\n<p><img src=\"https://aipaint.lihuanyu.com/66dca5b4a7937737a9bb6d68a1e29ff.png\" alt=\"微信接口安全测试\"></p>\n<p>当时的直觉是：服务端是自己写的，不是某个历史漏洞频发的开源系统；数据库只允许本机连接；接口也做了出入参校验，应该不会有太大问题。</p>\n<p>测试开始后，后台日志里出现了大量请求，但接口没有明显报错。真正发现异常，是在后台页面查看历史记录时，看到了几十条空白记录。进数据库一看，数据明显不是正常业务流程写进去的。</p>\n<p><img src=\"https://aipaint.lihuanyu.com/%E6%95%B0%E6%8D%AE%E5%BA%93%E8%A2%AB%E6%B3%A8%E5%85%A5%E5%90%8E%E7%9A%84%E5%BC%82%E5%B8%B8%E6%95%B0%E6%8D%AE.png\" alt=\"数据库异常数据\"></p>\n<p>这类问题有两个判断重点：</p>\n<ul>\n<li>是否只是写入了垃圾数据。</li>\n<li>是否存在越权读取、批量删除、数据泄露或权限扩大。</li>\n</ul>\n<p>实际现象只观察到异常写入，没有看到敏感数据泄露。但从安全视角看，只要攻击载荷能影响 SQL 语义，就不能按“小脏数据”处理。</p>\n<h2>从 PM2 日志定位入口</h2>\n<p>服务是用 PM2 部署的。实时日志可以用：</p>\n<pre><code class=\"language-bash\">pm2 logs\n</code></pre>\n<p>但排查历史请求时，更有用的是 PM2 写在磁盘上的日志文件。PM2 默认把日志保存到：</p>\n<pre><code class=\"language-text\">$HOME/.pm2/logs\n</code></pre>\n<p>可以把对应的 <code>out</code>、<code>error</code> 日志拉到本地，再用 <code>rg</code> 搜索关键字。比如：</p>\n<pre><code class=\"language-bash\">rg -n -i &quot;union select|sleep\\\\(|or 1=1|--|/\\\\*&quot; app-out.log\n</code></pre>\n<p>日志里当时能看到类似这样的请求痕迹：</p>\n<pre><code class=\"language-text\">用户 1676 请求历史绘图列表第 1&quot; union select 1,2-- 页\n</code></pre>\n<p>原始日志里还有一些终端颜色控制字符，看起来像乱码：</p>\n<p><img src=\"https://aipaint.lihuanyu.com/sql%E6%B3%A8%E5%85%A5%E6%97%A5%E5%BF%97%E5%88%86%E6%9E%90.png\" alt=\"SQL 注入日志分析\"></p>\n<p>这不是编码问题，而是 ANSI escape code。必要时可以先清理成纯文本再读：</p>\n<pre><code class=\"language-bash\">perl -pe 's/\\e\\[[0-9;]*[mK]//g' app-out.log &gt; app-out.clean.log\n</code></pre>\n<p>从日志看，攻击入口是“历史绘图列表”的分页参数。请求本来应该是第几页，结果页码位置被塞进了 SQL 注入 payload。</p>\n<h2>根因：把查询参数当成了可信数字</h2>\n<p>这类分页接口通常长这样：</p>\n<pre><code class=\"language-http\">GET /histories?page=1&amp;pageSize=20\n</code></pre>\n<p>问题出在服务端把 <code>page</code> 当作数字使用，但 HTTP 查询参数进入服务端时天然是字符串。只要没有强制转换和校验，<code>page</code> 就可能是任意字符串。</p>\n<p>危险写法通常有两种。</p>\n<p>第一种是直接拼 SQL：</p>\n<pre><code class=\"language-ts\">const sql = `\n  SELECT * FROM histories\n  WHERE user_id = ${userId}\n  ORDER BY created_at DESC\n  LIMIT ${(page - 1) * pageSize}, ${pageSize}\n`;\n</code></pre>\n<p>第二种是用了 ORM，但某些条件仍然用字符串拼接：</p>\n<pre><code class=\"language-ts\">queryBuilder\n  .where(`history.user_id = ${userId}`)\n  .take(pageSize)\n  .skip((page - 1) * pageSize);\n</code></pre>\n<p>这两种都把“不可信输入”放进了 SQL 结构里。只要输入没有被限制为真正的数字，就有被改变 SQL 语义的可能。</p>\n<p>修复不能只靠“把已经出现的 payload 过滤掉”。SQL 注入的关键不是某几个危险字符串，而是代码把用户输入当成 SQL 代码的一部分执行了。</p>\n<p>OWASP 对 SQL 注入防御的首要建议是参数化查询：SQL 结构先确定，用户输入作为参数绑定进去，让数据库始终能区分代码和数据。输入校验也很重要，但它不是参数化查询的替代品。</p>\n<h2>NestJS 里应该怎么校验参数</h2>\n<p>NestJS 提供了 Pipe 和 <code>ValidationPipe</code>，适合把“请求参数必须是什么类型”放在控制器边界上处理。</p>\n<p>对简单分页参数，可以直接使用内置 Pipe：</p>\n<pre><code class=\"language-ts\">import { DefaultValuePipe, ParseIntPipe, Query } from '@nestjs/common';\n\n@Get('histories')\nasync listHistories(\n  @Query('page', new DefaultValuePipe(1), ParseIntPipe) page: number,\n  @Query('pageSize', new DefaultValuePipe(20), ParseIntPipe) pageSize: number,\n) {\n  const safePage = Math.max(page, 1);\n  const safePageSize = Math.min(Math.max(pageSize, 1), 50);\n\n  return this.historyService.list({\n    page: safePage,\n    pageSize: safePageSize,\n  });\n}\n</code></pre>\n<p>如果参数更多，建议用 DTO：</p>\n<pre><code class=\"language-ts\">import { Type } from 'class-transformer';\nimport { IsInt, Max, Min } from 'class-validator';\n\nexport class ListHistoryQueryDto {\n  @Type(() =&gt; Number)\n  @IsInt()\n  @Min(1)\n  page = 1;\n\n  @Type(() =&gt; Number)\n  @IsInt()\n  @Min(1)\n  @Max(50)\n  pageSize = 20;\n}\n</code></pre>\n<p>全局启用 <code>ValidationPipe</code>：</p>\n<pre><code class=\"language-ts\">app.useGlobalPipes(\n  new ValidationPipe({\n    transform: true,\n    whitelist: true,\n    forbidNonWhitelisted: true,\n  }),\n);\n</code></pre>\n<p>几个细节很重要：</p>\n<ul>\n<li><code>transform: true</code> 让 DTO 有机会把 query string 转成数字。</li>\n<li><code>whitelist: true</code> 会移除 DTO 中未声明的字段。</li>\n<li><code>forbidNonWhitelisted: true</code> 会直接拒绝额外字段，而不是静默吞掉。</li>\n<li><code>@Type(() =&gt; Number)</code>、<code>@IsInt()</code>、<code>@Min()</code>、<code>@Max()</code> 要配合使用，不能只写 TypeScript 的 <code>number</code> 类型。</li>\n</ul>\n<p>TypeScript 类型只存在于编译期，HTTP 请求进来后仍然是字符串。服务端边界必须做运行时校验。</p>\n<h2>ORM 查询也要避免字符串拼接</h2>\n<p>使用 ORM 不代表天然免疫 SQL 注入。ORM 的安全性取决于是否使用了它的参数绑定能力。</p>\n<p>更稳妥的写法是使用 repository API：</p>\n<pre><code class=\"language-ts\">return this.historyRepository.find({\n  where: {\n    userId,\n  },\n  order: {\n    createdAt: 'DESC',\n  },\n  skip: (page - 1) * pageSize,\n  take: pageSize,\n});\n</code></pre>\n<p>如果必须用 QueryBuilder，条件也要参数化：</p>\n<pre><code class=\"language-ts\">return this.historyRepository\n  .createQueryBuilder('history')\n  .where('history.user_id = :userId', { userId })\n  .orderBy('history.created_at', 'DESC')\n  .skip((page - 1) * pageSize)\n  .take(pageSize)\n  .getMany();\n</code></pre>\n<p>不要这样写：</p>\n<pre><code class=\"language-ts\">.where(`history.user_id = ${userId}`)\n</code></pre>\n<p>更容易被忽略的是动态排序。字段名、表名、排序方向这类 SQL 结构通常不能用普通参数绑定解决，应该用 allow-list：</p>\n<pre><code class=\"language-ts\">const sortFields = {\n  createdAt: 'history.created_at',\n  id: 'history.id',\n} as const;\n\nconst sortDirections = {\n  asc: 'ASC',\n  desc: 'DESC',\n} as const;\n\nconst sortField = sortFields[query.sortBy] ?? sortFields.createdAt;\nconst sortDirection = sortDirections[query.order] ?? sortDirections.desc;\n\nqueryBuilder.orderBy(sortField, sortDirection);\n</code></pre>\n<p>这里不是把用户输入拼进去，而是把用户输入映射到服务端预先定义好的安全选项。</p>\n<h2>数据库约束是最后一道提醒</h2>\n<p>后台页面能看到空白记录，说明数据库层也缺少一些约束。</p>\n<p>业务上不应该为空的字段，数据库也应该明确表达：</p>\n<ul>\n<li><code>NOT NULL</code></li>\n<li>合理的 <code>VARCHAR</code> 长度</li>\n<li>枚举或状态字段约束</li>\n<li>外键或逻辑外键</li>\n<li><code>created_at</code>、<code>updated_at</code> 默认值</li>\n<li>必要的唯一索引</li>\n</ul>\n<p>数据库约束不能替代服务端校验，也不能防 SQL 注入。但当服务端漏掉某个边界时，数据库约束可以把问题从“悄悄写入脏数据”变成“写入失败并报警”。</p>\n<p>对于历史记录这类表，如果业务上必须有用户 ID、图片地址、状态、创建时间，就不应该允许空白行成功入库。</p>\n<h2>最小权限也不能省</h2>\n<p>SQL 注入的危害大小，和数据库账号权限直接相关。</p>\n<p>个人项目里很容易让应用使用一个权限很大的数据库账号，甚至能建表、删表、改结构。这样省事，但一旦出现注入，破坏半径会变大。</p>\n<p>更稳妥的做法是：</p>\n<ul>\n<li>应用运行账号只拥有业务所需的 <code>SELECT</code>、<code>INSERT</code>、<code>UPDATE</code>、<code>DELETE</code>。</li>\n<li>迁移账号和运行账号分开，建表改表不使用线上运行账号。</li>\n<li>不给应用账号 <code>DROP</code>、<code>ALTER</code>、全库管理权限。</li>\n<li>不同业务库使用不同账号，避免一个服务出问题影响所有数据。</li>\n</ul>\n<p>OWASP 也把 least privilege 作为 SQL 注入的纵深防御手段。它不能阻止漏洞出现，但能降低漏洞成功后的损失。</p>\n<h2>日志应该帮人定位，而不是只堆文本</h2>\n<p>问题能够定位，是因为日志里记录了关键请求。但原始日志仍然有几个不足：</p>\n<ul>\n<li>缺少统一 request id。</li>\n<li>参数、用户、接口路径没有结构化。</li>\n<li>日志里混有终端颜色控制字符。</li>\n<li>异常请求没有单独告警。</li>\n</ul>\n<p>服务端日志至少应该能回答这些问题：</p>\n<ul>\n<li>哪个用户或匿名标识发起了请求。</li>\n<li>请求路径和方法是什么。</li>\n<li>关键 query/body 参数是什么，敏感字段要脱敏。</li>\n<li>响应状态码和耗时是多少。</li>\n<li>异常堆栈和请求上下文如何关联。</li>\n</ul>\n<p>结构化日志比彩色文本更适合线上排查。即使不引入复杂日志系统，至少也可以让每条日志是 JSON，后续用 <code>rg</code>、<code>jq</code>、Loki、ELK 等工具都更容易处理。</p>\n<h2>安全测试不能只靠平台</h2>\n<p>微信的模拟攻击很有价值，帮助暴露了问题。但平台安全测试只能作为外部信号，不能替代服务端自身的安全工程。</p>\n<p>至少应该补几类测试：</p>\n<pre><code class=\"language-ts\">it('rejects non-numeric page query', async () =&gt; {\n  await request(app.getHttpServer())\n    .get('/histories?page=1%22%20union%20select%201,2--')\n    .expect(400);\n});\n\nit('limits pageSize', async () =&gt; {\n  await request(app.getHttpServer())\n    .get('/histories?page=1&amp;pageSize=10000')\n    .expect(400);\n});\n</code></pre>\n<p>还可以补服务层测试，确保分页参数经过 DTO 或 Pipe 后才进入查询逻辑；再补一条集成测试，确认异常请求不会写入任何业务记录。</p>\n<p>安全测试不需要一开始就很复杂。先把已经踩过的坑固化成测试，收益最高。</p>\n<h2>一份服务端安全复盘清单</h2>\n<p>类似问题处理后，可以按这张清单检查服务端项目：</p>\n<ol>\n<li>所有 <code>params</code>、<code>query</code>、<code>body</code> 是否都有运行时校验。</li>\n<li>分页参数是否是整数，并有最小值和最大值。</li>\n<li>动态排序字段是否使用 allow-list。</li>\n<li>ORM 查询是否使用参数绑定，是否还有 raw SQL 字符串拼接。</li>\n<li>数据库字段是否有必要的 <code>NOT NULL</code>、长度、索引和状态约束。</li>\n<li>线上应用数据库账号是否遵循最小权限。</li>\n<li>错误响应是否避免暴露 SQL、表名、堆栈和服务器路径。</li>\n<li>日志是否能按 request id 关联用户、接口、参数、状态码和异常。</li>\n<li>是否有针对已知攻击 payload 的 e2e 测试。</li>\n<li>是否有异常写入、异常错误率、异常 400/500 的监控。</li>\n</ol>\n<p>SQL 注入通常不是孤立问题。它背后往往同时有输入边界不清、查询写法不安全、数据库约束不足、日志不可观测等问题。</p>\n<h2>总结</h2>\n<p>这起事故的直接修复，是把分页参数强制转成数字并校验范围。但真正的复盘结论不止这一点。</p>\n<p>服务端安全要分层：</p>\n<ul>\n<li>Controller 边界用 Pipe/DTO 做运行时校验。</li>\n<li>查询层用参数化查询和 ORM 安全 API。</li>\n<li>动态 SQL 结构用 allow-list。</li>\n<li>数据库层用约束和最小权限降低损失。</li>\n<li>日志和测试负责让问题更早被发现、更容易复现。</li>\n</ul>\n<p>小程序平台的安全测试只是把问题推到了眼前。真正让系统变安全的，是把问题沉淀成代码约束、数据库约束、测试用例和排查流程。</p>\n<h2>扩展阅读</h2>\n<ul>\n<li><a href=\"https://cheatsheetseries.owasp.org/cheatsheets/SQL_Injection_Prevention_Cheat_Sheet.html\">OWASP: SQL Injection Prevention Cheat Sheet</a></li>\n<li><a href=\"https://cheatsheetseries.owasp.org/cheatsheets/Input_Validation_Cheat_Sheet.html\">OWASP: Input Validation Cheat Sheet</a></li>\n<li><a href=\"https://docs.nestjs.com/techniques/validation\">NestJS: Validation</a></li>\n<li><a href=\"https://docs.nestjs.com/pipes\">NestJS: Pipes</a></li>\n<li><a href=\"https://typeorm.io/docs/query-builder/select-query-builder\">TypeORM: Select using Query Builder</a></li>\n<li><a href=\"https://pm2.io/docs/runtime/guide/log-management/\">PM2: Log Management</a></li>\n</ul>\n","date_published":"2023-08-20T00:00:00.000Z","date_modified":"2026-05-04T00:00:00.000Z","tags":["SQL注入","NestJS","MySQL","服务端安全","PM2日志阅读"],"language":"zh"},{"id":"https://www.lihuanyu.com/posts/2023/%E5%B0%8F%E7%A8%8B%E5%BA%8F%E9%A1%B5%E9%9D%A2%E9%A1%B6%E9%83%A8%E7%9A%84%E7%A9%BA%E9%9A%99/","url":"https://www.lihuanyu.com/posts/2023/%E5%B0%8F%E7%A8%8B%E5%BA%8F%E9%A1%B5%E9%9D%A2%E9%A1%B6%E9%83%A8%E7%9A%84%E7%A9%BA%E9%9A%99/","title":"小程序页面顶部的空隙","summary":"解释小程序页面顶部可滚动空隙背后的 margin 塌陷问题，并比较空元素、BFC 和 overflow 方案的取舍。","content_html":"<blockquote>\n<p>TL;DR：移动端web页面顶上如果有空隙的话，可以对页面父元素用 padding 或者加空元素防止因 margin 塌陷造成的不正常滚动。</p>\n</blockquote>\n<h2>起源</h2>\n<p>强迫症同学有没有注意到，很多小程序的页面，明明不超过一页，但是却可以滚，但又只能滚一点点。</p>\n<p>比如这个：</p>\n<p><img src=\"/assets/legacy/_posts/%E5%B0%8F%E7%A8%8B%E5%BA%8F%E9%A1%B5%E9%9D%A2%E9%A1%B6%E9%83%A8%E7%9A%84%E7%A9%BA%E9%9A%99/margin-example.png\" alt=\"顶部空隙示例\"></p>\n<p>这种不符合预期的滚且只能滚一点点很难受，于是去探索到底是什么导致了这个小滚动的出现。</p>\n<p>用开发者工具去尝试找这个空隙是找不到的，但是能发现有这种表现的，无一不是最顶上的元素使用了 margin-top。</p>\n<h2>原因</h2>\n<p>在较长一段时间里，遇到这个问题的时候，会选择靠不在第一个元素上使用margin-top来规避这个问题的。</p>\n<p>有一天浏览微信开发者社区，发现也有同学有<a href=\"https://developers.weixin.qq.com/community/develop/doc/00080cc5040580a4e629d18f45ec00\">类似的问题</a>。</p>\n<p>微信小程序答疑同学表示这是： <code>margin-top 垂直方向塌陷导致的</code></p>\n<p>顺便给出了解决方案：</p>\n<pre><code class=\"language-html\">&lt;!--在第一个元素前加这样一个空元素--&gt;\n&lt;view style=&quot;content: ''; overflow: hidden;&quot;&gt;&lt;/view&gt;\n</code></pre>\n<p>试过了，很好用。但是交互强迫症满意了，代码强迫症犯了，页面最顶上要加这么个玩意儿？？？？</p>\n<p>这时可以注意到关键字，margin塌陷。搜索后才发现原来塌陷不光是曾经理解的两个元素间的margin会塌陷。元素套元素也会，看掘金的文章 - <a href=\"https://juejin.cn/post/6976272394247897101\">什么是margin塌陷及解决方案</a>。</p>\n<h2>优雅</h2>\n<p>所以其实关键是解决塌陷，加个空元素只是个手段。那有没有更优雅的手法？</p>\n<p>上面掘金的文章说可以用 <code>BFC</code> 来解决，写了好几种方法触发BFC：</p>\n<ol>\n<li>float 属性为 left / right</li>\n<li>overflow 为 hidden / scroll / auto</li>\n<li>position 为 absolute / fixed</li>\n<li>display 为 inline-block / table-cell / table-caption</li>\n</ol>\n<p>看起来 <code>overflow: auto</code> 是最安全的，给 page 加个这个能有什么坏处呢？</p>\n<p>在全局的 CSS 里给 page 元素加上这个样式，大功告成。</p>\n<h2>转折</h2>\n<p><code>overflow: auto</code> 并不是完全无害的， 加了这个会导致页面里的 <code>position: sticky</code> 失效。</p>\n<p><code>position: sticky</code> 要求父级元素不能有任何 <code>overflow:visible</code> 以外的overflow设置，否则没有粘滞效果。因为改变了滚动容器（即使没有出现滚动条）。</p>\n<p>更多细节可以看张鑫旭的文章 - <a href=\"https://www.zhangxinxu.com/wordpress/2018/12/css-position-sticky/\">position:sticky</a></p>\n<h2>结论</h2>\n<p>也许我们可以因地制宜地选择某些方法触发 BFC 来解决这个问题。但是如果需要选择，不如固定一种无害写法，虽然可能有点丑，但是能解决问题。</p>\n<p>遇到此问题时，直接在页面元素最前面加上 <code>&lt;view style=&quot;content: ''; overflow: hidden;&quot;&gt;&lt;/view&gt;</code> 来进行解决吧。</p>\n","date_published":"2023-02-19T00:00:00.000Z","tags":["前端","小程序"],"language":"zh"},{"id":"https://www.lihuanyu.com/posts/2022/%E7%94%BB%E5%9B%BE%E5%B7%A5%E5%85%B7-excalidraw/","url":"https://www.lihuanyu.com/posts/2022/%E7%94%BB%E5%9B%BE%E5%B7%A5%E5%85%B7-excalidraw/","title":"工程师画图工具：Excalidraw 和技术图表达","summary":"从 Excalidraw 这款手绘风格画图工具出发，讨论技术文章里的图应该如何帮助读者理解结构、流程、边界和取舍，而不只是让页面看起来热闹。","content_html":"<p>画图一直是我的弱项。</p>\n<p>后来慢慢发现，弱的可能不只是手，也包括脑子。画不出来的时候，常常不是工具不熟，而是自己还没想清楚。脑子里如果只有一团雾，换成什么画图工具，最后也只能画出一张雾的截图。</p>\n<p>不过工具仍然重要。</p>\n<p>好工具不能替人思考，但能降低表达的摩擦。尤其技术文章里，一张图如果画得清楚，读者能少走很多弯路。</p>\n<p>我后来比较喜欢用 <a href=\"https://excalidraw.com/\">Excalidraw</a>。</p>\n<h2>为什么是 Excalidraw</h2>\n<p>最早注意到 Excalidraw，是因为读 <a href=\"https://github.yanhaixiang.com/jest-tutorial/\">《Jest 实践指南》</a> 时，发现里面的配图很好看。</p>\n<p>比如这张：</p>\n<p><img src=\"/assets/legacy/_posts/%E7%94%BB%E5%9B%BE%E5%B7%A5%E5%85%B7-excalidraw/img.png\" alt=\"Jest实践指南中的配图\"></p>\n<p>它不是那种精确到像素的商务图，也不是满屏渐变和装饰的 PPT 图。线条有一点手绘感，形状很简单，重点很清楚。看起来不像工程制图，更像一个人在白板前把事情讲明白。</p>\n<p>后来搜了一下，发现工具就是 Excalidraw。</p>\n<p>它还有一个很对工程师胃口的地方：开源。代码在 <a href=\"https://github.com/excalidraw/excalidraw\">GitHub</a> 上，在线版本能直接用，也可以自己部署。如果在公司里对数据安全比较敏感，不想把图放到外部服务，私有化部署也有路径。</p>\n<p>当然，如果不想维护，也可以用它的增值服务。</p>\n<p>这些都不是最关键的。最关键的是，它的风格会逼人少装。</p>\n<p>手绘风格不太适合堆复杂视觉效果。你很难在 Excalidraw 里画出那种精致但空洞的咨询公司图。它更适合画关系、流程、层级、边界和变化。对技术文章来说，这刚好够用。</p>\n<h2>技术图不是插画</h2>\n<p>很多技术文章里的图，问题不是不好看，而是没有承担信息任务。</p>\n<p>有些图只是为了让页面不那么空，于是放一张看起来科技感十足的插画。读者看完以后，除了知道作者会配图，什么也没多明白。</p>\n<p>技术图应该先回答一个问题：</p>\n<p>这张图想替文字完成什么工作？</p>\n<p>如果只是为了装饰，它可有可无。如果能让读者更快理解一个结构、一个流程、一个状态变化、一个权衡关系，那它就有价值。</p>\n<p>我现在大概会把技术图分成几类：</p>\n<ol>\n<li>流程图：说明事情按什么顺序发生。</li>\n<li>架构图：说明模块之间怎么连接。</li>\n<li>状态图：说明对象会在哪些状态之间变化。</li>\n<li>对比图：说明两个方案差在哪里。</li>\n<li>分层图：说明一个系统有哪些边界。</li>\n<li>时间线：说明问题如何演进。</li>\n</ol>\n<p>画图前先选类型，比打开工具后乱拖形状重要得多。</p>\n<p>一篇文章如果讲的是“从请求到响应发生了什么”，就画流程；如果讲的是“前端、服务端、存储、队列之间的关系”，就画架构；如果讲的是“订单、任务、审核状态怎么变”，就画状态；如果讲的是“为什么不用 A 方案而用 B 方案”，就画对比。</p>\n<p>图的类型选错，后面越画越乱。</p>\n<h2>先写一句话，再画图</h2>\n<p>我觉得比较有用的办法，是画图前先写一句话。</p>\n<p>不是标题，而是这张图的结论。</p>\n<p>比如：</p>\n<blockquote>\n<p>AI 应用不是多一个网页，而是网页后面接了一条按需生产线。</p>\n</blockquote>\n<p>或者：</p>\n<blockquote>\n<p>前端 mock 的价值，是让页面在后端不完整时仍然能独立运转。</p>\n</blockquote>\n<p>这句话写清楚以后，图就不容易跑偏。</p>\n<p>所有元素都要为这句话服务。不能服务的，就删掉。技术图最怕什么都想放，最后变成一个缩小版系统全景。作者觉得完整，读者只觉得眼睛疼。</p>\n<p>一张图里最好只有一个主角。</p>\n<p>其他东西要么是背景，要么是支撑。读者第一眼应该知道从哪里看起，第二眼知道箭头往哪里走，第三眼能把图和正文里的论点对上。</p>\n<p>如果三眼之后还在找入口，这图大概率已经失败了。</p>\n<h2>Excalidraw 适合的练习方法</h2>\n<p>有了工具后，可以先临摹。</p>\n<p>我当时就照着《Jest 实践指南》里的图画了一张：</p>\n<p><img src=\"/assets/legacy/_posts/%E7%94%BB%E5%9B%BE%E5%B7%A5%E5%85%B7-excalidraw/img_1.png\" alt=\"临摹的jest的图例\"></p>\n<p>临摹不是抄袭发布，而是练手。看别人怎么分组、怎么留白、怎么用箭头、怎么控制文字长度。画几张以后就会发现，图好不好看，很多时候不在形状复杂，而在取舍。</p>\n<p>Excalidraw 里有几个习惯很有用：</p>\n<ol>\n<li>先用矩形和箭头搭骨架，不急着调颜色。</li>\n<li>一个区域只表达一层意思，不把概念堆成一坨。</li>\n<li>文字尽量短，能用名词就不用长句。</li>\n<li>同一类元素用同一种形状和颜色。</li>\n<li>箭头方向要稳定，少让读者绕路。</li>\n<li>留白要舍得，图不是越满越专业。</li>\n<li>最后再调字体、颜色和对齐。</li>\n</ol>\n<p>手绘风格还有一个好处：它降低了精确感。</p>\n<p>太正式的图，稍微没对齐就显得粗糙；手绘风格天然允许一点松动，读者也更容易把注意力放在关系上，而不是像素上。</p>\n<p>不过这不代表可以乱画。手绘感不是潦草，仍然要有清楚的层级和节奏。真拿笔画过就知道，手绘要好看其实很难。Excalidraw 是给了普通人一点白板表达能力，不是免除了表达能力。</p>\n<h2>图要和文章一起长出来</h2>\n<p>技术文章里的图，不应该是写完以后硬塞进去的装饰。</p>\n<p>更好的状态是，写到某个地方发现文字绕来绕去，读者可能要迷路，这时就该画图。图不是文章的花边，而是正文的一部分。</p>\n<p>有些地方适合用图替代长段解释：</p>\n<ol>\n<li>多个模块之间有调用关系。</li>\n<li>一个请求经过很多步骤。</li>\n<li>两个方案的差异需要并排看。</li>\n<li>某个概念本身是空间结构。</li>\n<li>某个问题的关键在边界，而不是细节。</li>\n</ol>\n<p>也有些地方不适合画图。</p>\n<p>比如观点判断、个人经历、价值取舍，文字反而更有力量。硬画一张“价值取舍模型图”，很容易把原本有血有肉的经验变成塑料流程。</p>\n<p>所以画图不是越多越好。</p>\n<p>一篇技术文章有一两张真正解决理解问题的图，比塞五张漂亮废话强。</p>\n<h2>工具只是最后一步</h2>\n<p>Excalidraw 是个好工具，但它不是画图能力本身。</p>\n<p>真正重要的是先把事情想清楚：主角是谁，关系是什么，变化在哪里，边界在哪里，读者为什么需要这张图。</p>\n<p>想清楚以后，工具只是把脑子里的结构搬出来。</p>\n<p>没想清楚时，工具越强，越容易把人带偏。模板、图标、素材、渐变、阴影，全都在招手。画着画着，原本想解释一个问题，最后做成了一张看起来很忙的海报。</p>\n<p>技术图最好的状态，是让读者忘了图本身。</p>\n<p>他只是顺着图看下去，突然明白了：哦，原来这里是这样连起来的。</p>\n<p>这就够了。</p>\n","date_published":"2022-09-26T00:00:00.000Z","date_modified":"2026-05-16T00:00:00.000Z","tags":["前端","画图","Excalidraw"],"language":"zh"},{"id":"https://www.lihuanyu.com/posts/2022/%E4%BA%A4%E4%BA%92%E7%9A%84%E6%84%8F%E4%B9%89/","url":"https://www.lihuanyu.com/posts/2022/%E4%BA%A4%E4%BA%92%E7%9A%84%E6%84%8F%E4%B9%89/","title":"交互的意义","summary":"从电饭煲预约煮粥这件小事出发，讨论好交互为什么不是功能更多，而是把用户需要理解的变量变少，让人用目标语言操作机器。","content_html":"<p>有些道理，读书时看见一百遍，也不如生活里摔一跤。</p>\n<p>交互就是这样。</p>\n<p>好的交互常常没有存在感。它不喊口号，不摆姿态，也不让人觉得产品经理很努力。用户只会觉得事情本该如此，手指点下去，东西就发生了。只有遇到糟糕的交互，人才突然意识到：原来一个按钮、一行文案、一种顺序，也能把人逼到墙角。</p>\n<p>我对这件事印象最深的一次，是电饭煲预约煮粥。</p>\n<h2>一口粥里的变量</h2>\n<p>有天晚上，我妈打电话问我，电饭煲怎么预约煮粥。</p>\n<p>为什么问我？因为我之前预约过，早上起来直接喝粥，看起来很熟练。</p>\n<p>可我一下子愣住了。</p>\n<p>我确实会预约，但我是在手机上操作的。打开米家，选煮粥，选预约，页面直接问我：想几点开饭？填一个时间，确认，就结束了。</p>\n<p>这不是我会用电饭煲，这是米家把问题翻译成了人话。</p>\n<p>我妈面对的不是这个界面。她面对的是一个传统电饭煲：按钮很多，模式很多，屏幕很小，说明书不知道塞到哪里去了。她不知道粥要煮多久，不知道预约时间是开始煮的时间，还是煮好的时间，也不知道设好时间后还要不要再按一次开始。</p>\n<p>于是一个很简单的目标，“明天早上喝粥”，被拆成了一堆机器变量：</p>\n<ol>\n<li>选择什么模式？</li>\n<li>煮粥需要多长时间？</li>\n<li>预约的是开始时间还是结束时间？</li>\n<li>现在离明早还有几个小时？</li>\n<li>设置完成后机器是否已经进入工作状态？</li>\n<li>如果设置错了，如何取消重来？</li>\n</ol>\n<p>这些变量对机器来说很自然，对人来说很荒唐。</p>\n<p>人想要的是粥，不是和电饭煲讨论调度算法。</p>\n<h2>好交互不是功能更多</h2>\n<p>很多产品谈“智能”，最后变成加联网、加 App、加按钮、加模式。东西确实多了，但人并不一定轻松。</p>\n<p>真正好的交互，不是把功能摆满，而是把用户需要理解的变量变少。</p>\n<p>传统电饭煲把“预约”交给用户自己计算。用户要知道现在几点、目标几点、煮多久、提前多久启动。米家那种方式则把问题倒过来：用户只说目标时间，机器自己算。</p>\n<p>这就是交互的价值。</p>\n<p>它不是把复杂性消灭了。复杂性还在，煮粥还是要时间，机器还是要执行步骤，传感器和程序还是要工作。只是复杂性被收回到了系统里，没有摊到用户脸上。</p>\n<p>好产品经常做的就是这件事：把机器语言翻译成人话。</p>\n<p>导航不问你“向东北方向行驶 1.7 公里后进入匝道编号 X”，它说前方右转；打车软件不让你研究司机调度，它问你从哪里到哪里；日历不要求你背时区规则，它只问几点开会。</p>\n<p>坏交互则相反。它把系统实现方式原封不动地交给用户，然后指望用户有耐心、有知识、有说明书、有空慢慢试。用户错了，它还会露出一种无辜的表情：功能都给你了，是你不会用。</p>\n<p>这很像某些早期后台系统。数据库里怎么存，页面就怎么展示；业务流程里有哪些状态，筛选项就堆多少个；接口需要什么字段，表单就让人填什么字段。开发者觉得忠实，用户觉得受刑。</p>\n<h2>互联网品牌的“降维打击”</h2>\n<p>互联网公司喜欢说“降维打击”。这个词被用烂了，但在一些硬件产品上，确实能看到类似现象。</p>\n<p>大量互联网品牌没有自己的工厂，只是出设计方案，找传统厂商生产，甚至直接套公模，贴自己的牌子。按制造能力看，它们未必更强。</p>\n<p>但我还是经常买这些所谓贴牌产品。</p>\n<p>一方面是外观设计通常在线，能和家里装修搭一点边。另一方面是软件体验常常更顺。联网本身不稀奇，加一个 Wi-Fi 模块不是什么天书；真正拉开差距的，是联网之后，用户到底是在控制设备，还是在继续伺候设备。</p>\n<p>家里以前主要有小米和美的两套智能家居。小米手机曾经辜负过我的信任，这笔账另算。但小米智能家居的交互，至少在我用过的那些设备里，确实经常比传统厂商顺。</p>\n<p>它不是每个功能都更强，而是更少让人背参数。</p>\n<p>对普通人来说，这比多几个模式重要得多。</p>\n<h2>少让用户证明自己聪明</h2>\n<p>很多糟糕交互有一个共同特点：它默认用户应该理解系统。</p>\n<p>用户应该知道预约是什么意思，应该知道粥多久煮好，应该知道按钮长按和短按的区别，应该知道图标代表什么，应该知道设置完还要确认。</p>\n<p>可用户为什么应该知道？</p>\n<p>人买电饭煲，是为了吃饭；买洗衣机，是为了衣服干净；买空调，是为了屋子舒服。用户不是来参加设备能力考试的。</p>\n<p>好的交互不是讨好用户，而是尊重现实：人的注意力有限，记忆有限，耐心有限，愿意花在设备上的理解成本更有限。</p>\n<p>所以很多设计问题，最后都可以落到一个朴素标准上：</p>\n<p>用户为了完成目标，需要理解多少不该由他理解的东西？</p>\n<p>需要理解得越多，交互越差。</p>\n<p>一个人如果只是想明早喝粥，却必须搞懂“预约是倒计时还是定时启动”，那就是系统把自己的麻烦推给了人。</p>\n<h2>回到那口粥</h2>\n<p>那天晚上，最后的结果很普通：我也没能在电话里把传统电饭煲的预约讲明白。</p>\n<p>不是我不会讲，而是这事不该这么讲。</p>\n<p>电话那头的人想听的是：“按这里，明早七点就能吃。”可机器给出的却是另一套话：“选择功能，设置小时，设置分钟，确认状态，注意指示灯。”</p>\n<p>两种语言之间隔着一条河。</p>\n<p>交互设计的意义，就是架桥。</p>\n<p>它不一定宏大，也不一定光鲜。很多时候只是把“开始煮”改成“几点吃”，把“模式参数”改成“我要做什么”，把一堆按钮变成一条清楚的路径。</p>\n<p>这事看起来小。可一个社会里的机器越来越多，App 越来越多，系统越来越多，小小的理解成本乘以无数次使用，就会变成很重的负担。</p>\n<p>好交互不是让人觉得机器聪明。</p>\n<p>好交互是让人不必证明自己聪明。</p>\n","date_published":"2022-09-25T00:00:00.000Z","date_modified":"2026-05-16T00:00:00.000Z","tags":["生活","随笔","交互"],"language":"zh"},{"id":"https://www.lihuanyu.com/posts/2022/CSS%E4%B9%8B%E5%9B%BE%E7%89%87%E4%B8%8B%E7%9A%84%E7%A9%BA%E9%9A%99%E4%B8%8E%E6%96%87%E6%9C%AC%E5%B1%85%E4%B8%AD/","url":"https://www.lihuanyu.com/posts/2022/CSS%E4%B9%8B%E5%9B%BE%E7%89%87%E4%B8%8B%E7%9A%84%E7%A9%BA%E9%9A%99%E4%B8%8E%E6%96%87%E6%9C%AC%E5%B1%85%E4%B8%AD/","title":"CSS之图片下的空隙与文本居中","summary":"解释图片底部空隙和文本垂直居中偏差背后的 CSS 行内元素、基线、line-height 与字体度量问题。","content_html":"<p>前端工程师应该都有遇到过，使用图片时会在下方有个小空隙。这个小空隙很难找到它是如何形成的，但是还好我们有搜索引擎，因此很容易会知道解决办法：</p>\n<p><img src=\"/assets/legacy/_posts/CSS%E4%B9%8B%E5%9B%BE%E7%89%87%E4%B8%8B%E7%9A%84%E7%A9%BA%E9%9A%99%E4%B8%8E%E6%96%87%E6%9C%AC%E5%B1%85%E4%B8%AD/search-in-google.png\" alt=\"搜索解决图片下空隙\"></p>\n<p>其中一种技巧非常简单有效：把字体设为0。很多时候可能就到此为止了。</p>\n<h2>引子</h2>\n<p>一直没有思考过，为什么图片下面会有一个空隙。直到逛知乎刷到尤雨溪的一篇回答，关于这个空隙的，非常通俗易懂，再进入到 <a href=\"https://www.zhihu.com/question/21558138\">对应的知乎问题</a> 看到各种大佬们的分析，很有意思，做个分享。</p>\n<p>里面有一篇译文，非常全面剖析了行内元素，CSS里文字的度量、行高（line-height）。如果想深入学习细节，建议<a href=\"https://zhuanlan.zhihu.com/p/25808995\">直接前往</a> 。(PS: 这篇文章我最认可的点是结论的第一条😀)</p>\n<h2>原因</h2>\n<p>本文的原因部分可以认为是我对这些大佬回答的理解，如果没看懂我写的，可以直接去原问题浏览更多解决方案及解释。</p>\n<p>图片作为行内元素，默认的对其方式的基线对其，基线是西文字体的概念，如图：</p>\n<p><img src=\"/assets/legacy/_posts/CSS%E4%B9%8B%E5%9B%BE%E7%89%87%E4%B8%8B%E7%9A%84%E7%A9%BA%E9%9A%99%E4%B8%8E%E6%96%87%E6%9C%AC%E5%B1%85%E4%B8%AD/what-is-base-line.png\" alt=\"什么是基线\"></p>\n<p>红线所示即为基线（baseline），注意看，文字的底线（bottom）和基线之间是有一段距离的，这个就是图片下有空隙的原因。</p>\n<p>再通俗一点：</p>\n<p><code>图片底部是基于文字基线的，而容器 div 的底部是低于基线的</code></p>\n<p>中文文字虽然没有基线的概念，但是也有留白区域，所以中文也有类似的问题。</p>\n<p>下面会有动手环节，提供相应代码方便有兴趣的同学可以快速自行验证。</p>\n<h2>解决方式</h2>\n<p>可以看到，既然问题是出现在文本的基线问题上。那么就围绕这一点来解决：</p>\n<p>img 设置 display:block<br>\nvertical-align:top/bottom/middle<br>\nfont-size设为 0 （注意对文本不能这么操作，手动狗头……）</p>\n<h2>衍生问题</h2>\n<p>难怪在移动端开发时，设计同学给出的设计稿，在还原后验收时，经常受到设计同学的灵魂拷问，这里怎么没居中。</p>\n<p>设计同学不知道的是，前端同学自己也很懵逼，我明明设置了line-height和高度一样，为什么就偏上偏下了？很可能就是系统下字体本身的问题。</p>\n<p>在比较小的按钮上效果会比较明显，考虑到大家的眼睛健康和视力问题，真诚地建议设计师不要追求一些“高级感”而把文字、按钮设计得过小。再结合这个问题，一定要小的话，别用边框了，也就不容易看出来。</p>\n<h2>动手试试</h2>\n<blockquote>\n<p>以下case使用codepen演示，无法加载的话可能需要科学上网。考虑到科学的门槛和速度，贴个图替代下。</p>\n</blockquote>\n<p>原始case：</p>\n<p><img src=\"https://aipaint.lihuanyu.com/%E5%B1%8F%E5%B9%95%E6%88%AA%E5%9B%BE%202023-03-12%20155733.png\" alt=\"图片的有缝、无缝情况\"></p>\n<p class=\"codepen\" data-height=\"300\" data-default-tab=\"html,result\" data-slug-hash=\"KKRVaWm\" data-user=\"sky-admin\" style=\"height: 300px; box-sizing: border-box; display: flex; align-items: center; justify-content: center; border: 2px solid; margin: 1em 0; padding: 1em;\">\n  <span>See the Pen <a href=\"https://codepen.io/sky-admin/pen/KKRVaWm\">\n  图片下的小空隙</a> by Huanyu Li (<a href=\"https://codepen.io/sky-admin\">@sky-admin</a>)\n  on <a href=\"https://codepen.io\">CodePen</a>.</span>\n</p>\n<script async src=\"https://cpwebassets.codepen.io/assets/embed/ei.js\"></script>\n<p>文字case：</p>\n<p><img src=\"https://aipaint.lihuanyu.com/%E5%B1%8F%E5%B9%95%E6%88%AA%E5%9B%BE%202023-03-12%20160029.png\" alt=\"文字与边框间的缝隙\"></p>\n<p class=\"codepen\" data-height=\"300\" data-default-tab=\"html,result\" data-slug-hash=\"qBYbRze\" data-user=\"sky-admin\" style=\"height: 300px; box-sizing: border-box; display: flex; align-items: center; justify-content: center; border: 2px solid; margin: 1em 0; padding: 1em;\">\n  <span>See the Pen <a href=\"https://codepen.io/sky-admin/pen/qBYbRze\">\n  Untitled</a> by Huanyu Li (<a href=\"https://codepen.io/sky-admin\">@sky-admin</a>)\n  on <a href=\"https://codepen.io\">CodePen</a>.</span>\n</p>\n<script async src=\"https://cpwebassets.codepen.io/assets/embed/ei.js\"></script>\n<p>做业务时不求甚解也许不是坏事，但是钻研深一分总有收获。</p>\n<h2>其他</h2>\n<p>img元素为什么默认是个行内元素呢？</p>\n<p>img元素是个比较早的元素，然后它本质上不是那张图，而是那个链接的占位符。于是作为占位符它默认就是个inline元素。</p>\n","date_published":"2022-03-06T00:00:00.000Z","tags":["CSS"],"language":"zh"},{"id":"https://www.lihuanyu.com/en/posts/2022/frontend-dependencies-lockfile-reproducible-builds/","url":"https://www.lihuanyu.com/en/posts/2022/frontend-dependencies-lockfile-reproducible-builds/","title":"Lockfiles Explained: npm ci and pnpm Frozen Installs","summary":"Commit the lockfile and use npm ci or pnpm --frozen-lockfile in CI so package manifests cannot silently resolve a different dependency tree.","content_html":"<p>A lockfile records the complete dependency tree selected for a project. Commit it, then use <code>npm ci</code> or <code>pnpm install --frozen-lockfile</code> in continuous integration (CI) and deployment. These commands fail when the package manifest and lockfile disagree instead of silently changing the dependency tree.</p>\n<p>Use normal install commands only when you intend to add, remove, or update dependencies:</p>\n<div class=\"table-scroll\"><table>\n<thead>\n<tr>\n<th>Task</th>\n<th>npm</th>\n<th>pnpm</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Install the committed dependency tree in CI</td>\n<td><code>npm ci</code></td>\n<td><code>pnpm install --frozen-lockfile</code></td>\n</tr>\n<tr>\n<td>Add a dependency</td>\n<td><code>npm install package_name</code></td>\n<td><code>pnpm add package_name</code></td>\n</tr>\n<tr>\n<td>Update dependencies intentionally</td>\n<td><code>npm update</code></td>\n<td><code>pnpm update</code></td>\n</tr>\n<tr>\n<td>Verify manifest and lockfile agreement</td>\n<td><code>npm ci</code></td>\n<td><code>pnpm install --frozen-lockfile</code></td>\n</tr>\n</tbody>\n</table>\n</div><p><a href=\"/posts/2022/%E5%89%8D%E7%AB%AF%E4%BE%9D%E8%B5%96%E4%B8%8E%E4%BF%A1%E4%BB%BB/\">Chinese version of this article</a></p>\n<p>The rest of this article explains what lockfiles guarantee, what they do not guarantee, and how to keep dependency changes reviewable.</p>\n<p>This old joke still works:</p>\n<p><img src=\"/assets/legacy/_posts/%E5%89%8D%E7%AB%AF%E4%BE%9D%E8%B5%96%E4%B8%8E%E4%BF%A1%E4%BB%BB/node_modules-black-hole.jpg\" alt=\"node_modules is heavier than a black hole\"></p>\n<p>Having many dependencies is not the problem by itself. The real question is this: when a project is built, are the installed dependencies exactly the same ones that were used during development, testing, and release validation?</p>\n<p>If the answer is no, a tiny feature change can ship with unexpected behavior caused by a dependency update that nobody reviewed. The goal of dependency management is not to reject third-party packages. The goal is to make dependency changes visible, controlled, and reversible.</p>\n<h2>package.json cannot pin the complete dependency tree</h2>\n<p>Frontend projects usually declare dependencies in <code>package.json</code>:</p>\n<pre><code class=\"language-json\">{\n  &quot;dependencies&quot;: {\n    &quot;some-package&quot;: &quot;^2.0.0&quot;\n  }\n}\n</code></pre>\n<p>Semantic Versioning splits a version into <code>X.Y.Z</code>:</p>\n<ul>\n<li><code>X</code> is the major version, usually used for incompatible changes.</li>\n<li><code>Y</code> is the minor version, usually used for backward-compatible features.</li>\n<li><code>Z</code> is the patch version, usually used for backward-compatible fixes.</li>\n</ul>\n<p>The full specification is here: <a href=\"https://semver.org/\">Semantic Versioning</a>.</p>\n<p><code>^2.0.0</code> does not mean “always install 2.0.0”. It means npm can install a compatible version in the allowed range. In practice, that may be <code>2.1.0</code> or <code>2.3.4</code>, as long as it stays within the compatible <code>2.x</code> range. <code>~2.0.0</code> is more conservative and usually allows patch-level changes.</p>\n<p>This design is reasonable. Patch releases fix bugs, minor releases add capabilities, and projects can benefit from maintenance automatically. But it relies on one assumption: package maintainers publish compatible releases and do not introduce security or quality problems in later versions.</p>\n<p>That assumption is not always true. Maintainers can publish by mistake, underestimate a breaking change, or intentionally publish destructive code. The colors.js/faker.js incident is a well-known example: a maintainer released versions with disruptive behavior, and many downstream dependency chains were affected. Supply chain incidents like that cannot be solved by trusting version numbers alone.</p>\n<p>There is another problem: pinning only direct dependencies is still not enough.</p>\n<h2>Transitive dependencies are part of every build</h2>\n<p>Writing exact versions in <code>package.json</code> looks safer:</p>\n<pre><code class=\"language-json\">{\n  &quot;dependencies&quot;: {\n    &quot;some-package&quot;: &quot;2.0.0&quot;\n  }\n}\n</code></pre>\n<p>That only pins dependencies declared directly by the project. A frontend package often depends on other packages, and those packages depend on more packages. Open the <code>node_modules</code> directory of a real project and the structure often looks like this:</p>\n<p><img src=\"/assets/legacy/_posts/%E5%89%8D%E7%AB%AF%E4%BE%9D%E8%B5%96%E4%B8%8E%E4%BF%A1%E4%BB%BB/deps-example.png\" alt=\"deep dependency tree example\"></p>\n<p>Even if every direct dependency avoids <code>^</code> and <code>~</code>, the dependencies of those dependencies may still use ranges. What actually participates in a build is the whole dependency tree, not just the few lines visible in <code>package.json</code>.</p>\n<p>So the thing worth locking is not “a few direct versions”. It is the complete dependency tree resolved during a known install.</p>\n<h2>Lockfiles preserve dependency resolution</h2>\n<p>npm’s <code>package-lock.json</code>, pnpm’s <code>pnpm-lock.yaml</code>, and Yarn’s <code>yarn.lock</code> solve the same core problem: they record the full dependency resolution result of an install.</p>\n<p>For example, <code>package-lock.json</code> records:</p>\n<ul>\n<li>The exact version of each package in the dependency tree.</li>\n<li>Where each package was resolved from, such as a registry tarball or a git commit.</li>\n<li>Integrity information for package contents.</li>\n<li>The relationship between dependencies.</li>\n</ul>\n<p>npm’s documentation also states that <code>package-lock.json</code> is intended to describe a dependency tree so teammates, deployments, and CI can install the same tree. That is why lockfiles should be committed to source control.</p>\n<p>With a lockfile, a project moves from “resolve dependencies from version ranges every time” to “reproduce a previously resolved dependency tree”. This is the foundation for reproducible builds.</p>\n<p>Here, “reproducible build” does not mean a fully cryptographically proven build process. In this context, it means meeting several practical engineering expectations:</p>\n<ul>\n<li>The same commit installs the same dependencies on different machines.</li>\n<li>CI, testing, and deployment use the same dependency resolution result.</li>\n<li>Dependency changes appear as lockfile diffs during code review.</li>\n<li>A build failure can be reproduced by checking out a specific Git commit.</li>\n</ul>\n<h2>Choose npm install or npm ci by intent</h2>\n<p><code>npm install</code> and <code>npm ci</code> both install dependencies, but they are meant for different situations.</p>\n<p><code>npm install</code> is the command for normal dependency maintenance. It reads <code>package.json</code> and <code>package-lock.json</code>. If the versions in the lockfile still satisfy the ranges in <code>package.json</code>, npm can keep using the locked versions. If not, npm resolves dependencies again and updates the lockfile.</p>\n<p>So <code>npm install</code> fits these cases:</p>\n<ul>\n<li>Initial dependency installation.</li>\n<li>Adding a dependency.</li>\n<li>Removing a dependency.</li>\n<li>Upgrading a dependency.</li>\n<li>Changing a dependency range.</li>\n</ul>\n<p><code>npm ci</code> is better suited for automated environments. According to npm’s documentation, it is designed for test platforms, continuous integration, and deployment. Its key behaviors are:</p>\n<ul>\n<li>It requires an existing <code>package-lock.json</code> or <code>npm-shrinkwrap.json</code>.</li>\n<li>If the lockfile does not match <code>package.json</code>, it fails instead of updating the lockfile.</li>\n<li>It removes the existing <code>node_modules</code> directory before installation.</li>\n<li>It does not write to <code>package.json</code> or the lockfile. The install is frozen.</li>\n</ul>\n<p>That is exactly what CI/CD needs. If dependency descriptions are inconsistent, the build should expose the problem instead of silently generating a new dependency graph on the build machine.</p>\n<p>For npm projects, a basic workflow can be:</p>\n<pre><code class=\"language-bash\"># When intentionally adding or upgrading a dependency\nnpm install some-package\n\n# After cloning, switching branches, reinstalling dependencies, debugging, CI, and deployment\nnpm ci\n</code></pre>\n<p>If the original <code>package-lock.json</code> was generated with npm configuration that affects the dependency tree, such as <code>legacy-peer-deps</code> or <code>install-links</code>, those options should be saved in a project-level <code>.npmrc</code> and committed to the repository. Otherwise, <code>npm ci</code> may fail in another environment.</p>\n<h2>Use frozen installs with pnpm</h2>\n<p>pnpm follows the same idea with different commands.</p>\n<p>pnpm projects commit <code>pnpm-lock.yaml</code>. In CI environments, if a lockfile exists but would need to be updated, <code>pnpm install</code> fails by default. The explicit command is:</p>\n<pre><code class=\"language-bash\">pnpm install --frozen-lockfile\n</code></pre>\n<p>The intent is direct: do not update the lockfile; if the lockfile and manifest are inconsistent, fail the install.</p>\n<p>For pnpm projects, a basic workflow can be:</p>\n<pre><code class=\"language-bash\"># When intentionally adding or upgrading a dependency\npnpm add some-package\n\n# After cloning, switching branches, reinstalling dependencies, debugging, CI, and deployment\npnpm install --frozen-lockfile\n</code></pre>\n<p>For monorepos, workspace scope matters too. A dependency change can affect multiple packages, and lockfile diffs can become larger. In that kind of project, dependency upgrades should be separated from normal feature changes, reviewed independently, and verified explicitly.</p>\n<h2>Pin the package manager and Node.js version</h2>\n<p>A lockfile pins the dependency tree, but different package managers and different major versions can have different resolution behavior and lockfile formats. If some developers use npm, others use pnpm, or a project is edited with very different pnpm versions, the lockfile can change for reasons unrelated to the actual application.</p>\n<p>A project should pin this information:</p>\n<pre><code class=\"language-json\">{\n  &quot;packageManager&quot;: &quot;pnpm@10.10.0&quot;,\n  &quot;engines&quot;: {\n    &quot;node&quot;: &quot;&gt;=24 &lt;25&quot;\n  }\n}\n</code></pre>\n<p><code>packageManager</code> tells tooling which package manager and version the project expects. Used with Corepack or a team convention, it reduces unnecessary lockfile churn caused by local tooling differences.</p>\n<p>Node.js should also be pinned. A project can use <code>.nvmrc</code>, Volta, asdf, mise, or CI configuration for that. The specific tool is less important than the goal: development machines, CI, and deployment should share the same runtime assumptions.</p>\n<h2>Make dependency updates explicit</h2>\n<p>Reproducible builds do not mean never upgrading dependencies. Refusing upgrades forever creates another problem: security fixes are missed, ecosystem compatibility drifts, and the eventual upgrade becomes much more expensive.</p>\n<p>A healthier approach is to make dependency updates explicit:</p>\n<ol>\n<li>Avoid mixing incidental dependency upgrades into normal feature work.</li>\n<li>When adding or upgrading dependencies, make a separate commit so <code>package.json</code> and lockfile diffs are easy to review.</li>\n<li>Always use frozen installs in CI and deployment.</li>\n<li>Do dependency maintenance on a schedule, such as every two weeks or every month.</li>\n<li>After dependency upgrades, run the full test and build process, and do manual regression testing when necessary.</li>\n</ol>\n<p>This keeps dependency changes out of unrelated business diffs. During review, it becomes clear which package changed, which transitive dependencies were affected, whether new install scripts appeared, and whether any dependency source changed from a registry package to a git, tarball, or URL source.</p>\n<p>A lockfile diff does not need to be read line by line, but several signals are worth checking:</p>\n<ul>\n<li>Whether unfamiliar high-risk packages appeared.</li>\n<li>Whether many new transitive dependencies were added.</li>\n<li>Whether a package source changed from registry to git, tarball, or URL.</li>\n<li>Whether new <code>postinstall</code>, <code>install</code>, or <code>preinstall</code> scripts appeared.</li>\n<li>Whether any dependency crossed a major version boundary.</li>\n<li>Whether the package manager version or lockfile version changed.</li>\n</ul>\n<p>Tools such as <code>npm audit</code>, <code>pnpm audit</code>, and GitHub Dependabot can provide useful security signals, but <code>audit fix --force</code> should not be treated as a harmless automatic cleanup. It may introduce major upgrades or behavior changes. Security fixes still need to go through the normal test and release process.</p>\n<h2>Do not commit node_modules</h2>\n<p>Some people suggest committing <code>node_modules</code> to the repository. The motivation is understandable: if the install step is risky, put the install result under version control too.</p>\n<p>For most frontend application projects, that is not a good default:</p>\n<ul>\n<li>Repository size grows dramatically.</li>\n<li>Diffs become difficult to review.</li>\n<li>Native dependencies can behave differently across operating systems and CPU architectures.</li>\n<li>Install scripts, generated files, and symlinks are not always a good fit for Git.</li>\n<li>Daily development and code hosting become slower and less pleasant.</li>\n</ul>\n<p>Large projects such as Chrome have their own engineering constraints and infrastructure. Their choices should not be copied directly into normal web projects. The more common practice is still: commit the lockfile, do not commit <code>node_modules</code>, and use frozen installs in CI and deployment to reproduce dependencies.</p>\n<h2>Use this lockfile workflow</h2>\n<p>The practical checklist is:</p>\n<ul>\n<li>Commit <code>package-lock.json</code>, <code>pnpm-lock.yaml</code>, or <code>yarn.lock</code>.</li>\n<li>Do not commit <code>node_modules</code>.</li>\n<li>Use <code>npm ci</code> or <code>pnpm install --frozen-lockfile</code> in CI, testing, and deployment.</li>\n<li>Prefer frozen installs after cloning, switching branches, or reinstalling dependencies locally.</li>\n<li>Use <code>npm install</code>, <code>pnpm add</code>, <code>pnpm update</code>, and similar commands only when intentionally adding, removing, or upgrading dependencies.</li>\n<li>Keep dependency changes in separate commits when possible.</li>\n<li>Pin Node.js and the package manager version to avoid lockfile churn caused by tooling differences.</li>\n<li>Commit project-level npm or pnpm configuration, especially settings that affect dependency resolution.</li>\n<li>Use audit and Dependabot-style tools for security signals, but send fixes through normal testing and release flow.</li>\n<li>Think before adding a dependency. A smaller dependency surface is easier to maintain.</li>\n</ul>\n<p>Frontend dependency management cannot eliminate every supply chain risk. But it can turn risk from “something random that happened during a build” into “a visible change that appeared during code review”. That is the main value of lockfiles and frozen installs.</p>\n<h2>Further reading</h2>\n<ul>\n<li><a href=\"https://docs.npmjs.com/cli/v11/configuring-npm/package-lock-json/\">package-lock.json | npm Docs</a></li>\n<li><a href=\"https://docs.npmjs.com/cli/v11/commands/npm-ci/\">npm ci | npm Docs</a></li>\n<li><a href=\"https://pnpm.io/cli/install\">pnpm install | pnpm</a></li>\n<li><a href=\"https://nodejs.org/api/corepack.html\">Corepack | Node.js Documentation</a></li>\n<li><a href=\"https://semver.org/\">Semantic Versioning 2.0.0</a></li>\n</ul>\n","date_published":"2022-01-11T00:00:00.000Z","date_modified":"2026-08-01T00:00:00.000Z","tags":["Frontend","JavaScript","npm","pnpm","Lockfile","Reproducible Builds"],"language":"en"},{"id":"https://www.lihuanyu.com/posts/2022/%E5%89%8D%E7%AB%AF%E4%BE%9D%E8%B5%96%E4%B8%8E%E4%BF%A1%E4%BB%BB/","url":"https://www.lihuanyu.com/posts/2022/%E5%89%8D%E7%AB%AF%E4%BE%9D%E8%B5%96%E4%B8%8E%E4%BF%A1%E4%BB%BB/","title":"前端依赖、lockfile 与可信构建","summary":"从 npm 依赖版本、传递依赖、lockfile、npm ci 和 pnpm frozen install 出发，整理前端项目如何获得更稳定、可复现的构建结果。","content_html":"<p>前端项目很少真正“只写自己的代码”。从构建工具、框架、组件库，到日期处理、请求封装、样式处理，一个项目背后通常站着一整棵依赖树。</p>\n<p>组件也不一定只能通过 npm 包进入项目。<a href=\"/posts/2024/shadcn-ui%E7%BB%84%E4%BB%B6%E5%BA%93/\">shadcn/ui 的源码分发方式</a>会把组件文件直接加入仓库，把升级问题从“更新包版本”变成“审查并合并上游源码差异”。两种方式的所有权不同，但都需要让变化可见。</p>\n<p>npm 生态的好处非常明显：发包门槛低，社区包丰富，很多能力不用重复造轮子。代价也同样明显：依赖链很深，包质量参差不齐，越是现代化的项目，越容易拥有一个庞大的 <code>node_modules</code>。</p>\n<p><a href=\"/en/posts/2022/frontend-dependencies-lockfile-reproducible-builds/\">English version: Frontend Dependencies, Lockfiles, and Reproducible Builds</a></p>\n<p>这张老图仍然很传神：</p>\n<p><img src=\"/assets/legacy/_posts/%E5%89%8D%E7%AB%AF%E4%BE%9D%E8%B5%96%E4%B8%8E%E4%BF%A1%E4%BB%BB/node_modules-black-hole.jpg\" alt=\"比黑洞还重的是 node_modules\"></p>\n<p>依赖多不是原罪。真正的问题是：构建时安装到的依赖，是否就是开发、测试和发布时验证过的那一份？</p>\n<p>如果答案是否定的，一个很小的需求改动，也可能因为某个依赖的变化让线上产物出现非预期行为。前端依赖管理的核心，不是拒绝第三方包，而是让依赖变化变得可见、可控、可回滚。</p>\n<h2>package.json 不够</h2>\n<p>前端项目通常在 <code>package.json</code> 里声明依赖：</p>\n<pre><code class=\"language-json\">{\n  &quot;dependencies&quot;: {\n    &quot;some-package&quot;: &quot;^2.0.0&quot;\n  }\n}\n</code></pre>\n<p>语义化版本约定把版本号拆成 <code>X.Y.Z</code>：</p>\n<ul>\n<li><code>X</code> 是主版本号，通常用于不兼容变更。</li>\n<li><code>Y</code> 是次版本号，通常用于向下兼容的新能力。</li>\n<li><code>Z</code> 是修订号，通常用于向下兼容的问题修复。</li>\n</ul>\n<p>完整规范可以看 <a href=\"https://semver.org/lang/zh-CN/\">Semantic Versioning</a>。</p>\n<p><code>^2.0.0</code> 的含义不是“永远安装 2.0.0”，而是在兼容范围内安装满足条件的版本。实际安装时，可能拿到 <code>2.1.0</code>、<code>2.3.4</code>，只要仍在 <code>2.x</code> 的范围内即可。<code>~2.0.0</code> 更保守一些，通常只允许修订号变化。</p>\n<p>这种设计本身合理：补丁版本修 bug、次版本加能力，项目可以自动获得维护收益。但它依赖一个前提：包作者正确遵守语义化版本，且后续发布版本没有安全或质量问题。</p>\n<p>现实里这个前提并不总是成立。维护者可能误发、可能低估 breaking change，也可能主动发布破坏性代码。colors.js/faker.js 事件就是典型例子：维护者发布带破坏行为的版本后，大量依赖链受到影响。类似的供应链事故很难完全靠“相信版本号”解决。</p>\n<p>更麻烦的是，固定直接依赖版本也不够。</p>\n<h2>传递依赖才是深水区</h2>\n<p>把 <code>package.json</code> 里的依赖都写成精确版本，看起来能减少波动：</p>\n<pre><code class=\"language-json\">{\n  &quot;dependencies&quot;: {\n    &quot;some-package&quot;: &quot;2.0.0&quot;\n  }\n}\n</code></pre>\n<p>这只能锁住项目直接声明的依赖。一个前端包往往还会依赖其他包，其他包再继续依赖更多包。随便打开一个项目的 <code>node_modules</code>，经常能看到这样的结构：</p>\n<p><img src=\"/assets/legacy/_posts/%E5%89%8D%E7%AB%AF%E4%BE%9D%E8%B5%96%E4%B8%8E%E4%BF%A1%E4%BB%BB/deps-example.png\" alt=\"深层依赖示例\"></p>\n<p>直接依赖不带 <code>^</code> 和 <code>~</code>，并不代表它的依赖也全部固定。真正参与构建的是完整依赖树，而不是 <code>package.json</code> 里能看到的那几行。</p>\n<p>因此，前端项目需要锁住的不是“几个直接依赖的版本号”，而是“某一次安装解析出的整棵依赖树”。</p>\n<h2>lockfile 解决什么</h2>\n<p>npm 的 <code>package-lock.json</code>、pnpm 的 <code>pnpm-lock.yaml</code>、Yarn 的 <code>yarn.lock</code>，解决的是同一个核心问题：记录一次安装得到的完整依赖解析结果。</p>\n<p>以 <code>package-lock.json</code> 为例，它会记录：</p>\n<ul>\n<li>依赖树里每个包的具体版本。</li>\n<li>包从哪里解析而来，例如 registry tarball 或 git commit。</li>\n<li>包内容的完整性校验信息，例如 <code>integrity</code>。</li>\n<li>依赖之间的关系。</li>\n</ul>\n<p>npm 官方文档也明确说明，<code>package-lock.json</code> 用来描述一份依赖树，使团队成员、部署环境和 CI 可以安装到完全相同的依赖。这也是 lockfile 应该提交到源码仓库的原因。</p>\n<p>有了 lockfile，项目就从“每次按版本范围重新解析依赖”，变成“优先复现上次已经解析过的依赖树”。这一步是可信构建的基础。</p>\n<p>这里的“可信构建”不是说构建过程已经具备密码学意义上的完全可证明性，而是指至少满足几个工程要求：</p>\n<ul>\n<li>同一个提交在不同机器上安装到相同依赖。</li>\n<li>CI、测试、部署使用同一份依赖解析结果。</li>\n<li>依赖变化以 lockfile diff 的形式进入代码审查。</li>\n<li>构建失败时可以回到某个 Git 提交复现现场。</li>\n</ul>\n<h2>npm install 与 npm ci</h2>\n<p><code>npm install</code> 和 <code>npm ci</code> 都能安装依赖，但定位不同。</p>\n<p><code>npm install</code> 是日常维护依赖的命令。它会读取 <code>package.json</code> 和 <code>package-lock.json</code>，如果 lockfile 里的版本仍满足 <code>package.json</code> 的版本范围，npm 会继续使用 lockfile 里的具体版本；如果不满足，npm 会重新解析并更新 lockfile。</p>\n<p>所以 <code>npm install</code> 适合这些场景：</p>\n<ul>\n<li>初始化项目依赖。</li>\n<li>新增依赖。</li>\n<li>删除依赖。</li>\n<li>升级依赖。</li>\n<li>修改依赖版本范围。</li>\n</ul>\n<p><code>npm ci</code> 更适合自动化环境。根据 npm 文档，它面向测试平台、持续集成和部署等场景。它的关键行为是：</p>\n<ul>\n<li>必须存在 <code>package-lock.json</code> 或 <code>npm-shrinkwrap.json</code>。</li>\n<li>如果 lockfile 与 <code>package.json</code> 不匹配，直接失败，而不是自动更新 lockfile。</li>\n<li>安装前会清理已有的 <code>node_modules</code>。</li>\n<li>不会写入 <code>package.json</code> 或 lockfile，安装过程是 frozen 的。</li>\n</ul>\n<p>这正是 CI/CD 需要的行为：如果依赖描述不一致，应当暴露问题，而不是在构建机器上悄悄生成一份新的依赖图。</p>\n<p>对于 npm 项目，一个基础流程可以这样定：</p>\n<pre><code class=\"language-bash\"># 开发者明确要新增或升级依赖时\nnpm install some-package\n\n# 刚拉代码、切分支、重装依赖、排查问题、CI 和部署时\nnpm ci\n</code></pre>\n<p>如果生成 <code>package-lock.json</code> 时使用过会影响依赖树形状的 npm 配置，例如 <code>legacy-peer-deps</code> 或 <code>install-links</code>，这些配置也应该沉淀到项目级 <code>.npmrc</code> 并提交到仓库，否则 <code>npm ci</code> 在其他环境可能安装失败。</p>\n<h2>pnpm 项目怎么做</h2>\n<p>pnpm 的思路类似，但命令不同。</p>\n<p>pnpm 项目提交的是 <code>pnpm-lock.yaml</code>。在 CI 环境里，如果存在 lockfile 但它需要更新，<code>pnpm install</code> 默认会失败；显式写法是：</p>\n<pre><code class=\"language-bash\">pnpm install --frozen-lockfile\n</code></pre>\n<p>这个命令表达得更直接：不更新 lockfile；如果 lockfile 与 manifest 不一致，就让安装失败。</p>\n<p>所以 pnpm 项目的基础流程可以这样定：</p>\n<pre><code class=\"language-bash\"># 开发者明确要新增或升级依赖时\npnpm add some-package\n\n# 刚拉代码、切分支、重装依赖、排查问题、CI 和部署时\npnpm install --frozen-lockfile\n</code></pre>\n<p>如果项目使用 monorepo，还要注意 workspace 范围。依赖变化可能影响多个包，lockfile diff 也会更大。越是这种场景，越应该把依赖升级从普通业务改动里拆出来，单独提交、单独验证。</p>\n<h2>包管理器版本也要固定</h2>\n<p>lockfile 锁住的是依赖树，但不同包管理器、不同主版本的解析算法和 lockfile 格式也可能不同。团队里有人用 npm，有人用 pnpm，或者同一个项目里 pnpm 版本跨度太大，都可能让 lockfile 产生不必要的变化。</p>\n<p>项目最好同时固定这些信息：</p>\n<pre><code class=\"language-json\">{\n  &quot;packageManager&quot;: &quot;pnpm@10.10.0&quot;,\n  &quot;engines&quot;: {\n    &quot;node&quot;: &quot;&gt;=24 &lt;25&quot;\n  }\n}\n</code></pre>\n<p><code>packageManager</code> 能让工具知道这个项目期望使用哪个包管理器及版本。配合 Corepack 或团队约定，可以减少“我本地 pnpm 版本不一样所以 lockfile 变了”的问题。</p>\n<p>Node 版本也应该固定。可以用 <code>.nvmrc</code>、Volta、asdf、mise 或 CI 配置来约束。核心目标不是追求某个工具，而是让开发机、CI、部署机使用同一组运行时前提。</p>\n<h2>依赖更新应该是一个显式动作</h2>\n<p>可信构建不是永远不升级依赖。长期不升级会带来另一个问题：漏洞修复拿不到，生态适配不上，最终一次性升级成本更高。</p>\n<p>更合理的做法是把依赖更新变成显式动作：</p>\n<ol>\n<li>普通业务开发尽量不要顺手升级依赖。</li>\n<li>新增或升级依赖时单独提交，让 <code>package.json</code> 和 lockfile diff 容易审查。</li>\n<li>CI 和部署始终使用 frozen install。</li>\n<li>定期做依赖维护，例如每两周或每月集中处理一次。</li>\n<li>依赖升级后跑完整测试和构建，必要时补一次人工回归。</li>\n</ol>\n<p>这样做的好处是，依赖变化不会混在业务 diff 里。代码审查时可以清楚看到：升级的是哪个包、带来了哪些传递依赖变化、有没有新的 install script、有没有替换 registry 或 git 来源。</p>\n<p>lockfile diff 不需要逐行读完，但几个信号值得关注：</p>\n<ul>\n<li>是否出现陌生的高风险包。</li>\n<li>是否新增大量传递依赖。</li>\n<li>是否从 registry 包变成 git/tarball/url 来源。</li>\n<li>是否出现新的 <code>postinstall</code>、<code>install</code>、<code>preinstall</code> 脚本。</li>\n<li>是否有跨主版本升级。</li>\n<li>是否改动了包管理器版本或 lockfileVersion。</li>\n</ul>\n<p><code>npm audit</code>、<code>pnpm audit</code>、GitHub Dependabot 这类工具可以提供安全信号，但不适合无脑 <code>audit fix --force</code>。自动修复可能跨主版本升级，也可能引入新的行为变化。安全修复仍然要进入正常的测试和发布流程。</p>\n<h2>不要提交 node_modules</h2>\n<p>偶尔会有人建议把 <code>node_modules</code> 一起提交到仓库。这个思路的动机可以理解：既然担心安装阶段变化，那就把安装结果也纳入版本控制。</p>\n<p>但对绝大多数前端业务项目来说，这不是一个好默认值：</p>\n<ul>\n<li>仓库体积会急剧膨胀。</li>\n<li>diff 很难审查。</li>\n<li>跨系统、跨 CPU 架构、原生依赖会更麻烦。</li>\n<li>安装脚本、构建产物、软链接等细节不一定适合直接进 Git。</li>\n<li>团队日常开发和代码托管体验都会变差。</li>\n</ul>\n<p>Chrome 这类超大型项目有自己的工程背景和基础设施，不能直接套到普通 Web 项目上。更常规的做法仍然是：提交 lockfile，不提交 <code>node_modules</code>，在 CI/部署阶段用 frozen install 复现依赖。</p>\n<h2>推荐实践</h2>\n<p>整理成一份可执行的清单：</p>\n<ul>\n<li>提交 <code>package-lock.json</code>、<code>pnpm-lock.yaml</code> 或 <code>yarn.lock</code>。</li>\n<li>不提交 <code>node_modules</code>。</li>\n<li>CI、测试、部署使用 <code>npm ci</code> 或 <code>pnpm install --frozen-lockfile</code>。</li>\n<li>开发者刚拉代码、切分支、重装依赖时，也优先用 frozen install。</li>\n<li>只有新增、删除、升级依赖时，才使用 <code>npm install</code>、<code>pnpm add</code>、<code>pnpm update</code> 等会修改 lockfile 的命令。</li>\n<li>依赖变更尽量单独提交，便于审查和回滚。</li>\n<li>固定 Node 和包管理器版本，避免工具版本差异导致 lockfile 抖动。</li>\n<li>项目级 npm/pnpm 配置要提交到仓库，特别是会影响依赖解析的配置。</li>\n<li>使用 audit、Dependabot 等工具获取安全信号，但把修复纳入正常测试发布流程。</li>\n<li>添加依赖前先判断是否真的需要，越小的依赖面越容易维护。</li>\n</ul>\n<p>前端依赖管理不可能消除所有供应链风险，但可以把风险从“构建时随机发生”变成“代码审查时显式出现”。这就是 lockfile 和 frozen install 最重要的价值。</p>\n<h2>扩展阅读</h2>\n<ul>\n<li><a href=\"https://docs.npmjs.com/cli/v11/configuring-npm/package-lock-json/\">package-lock.json | npm Docs</a></li>\n<li><a href=\"https://docs.npmjs.com/cli/v11/commands/npm-ci/\">npm ci | npm Docs</a></li>\n<li><a href=\"https://pnpm.io/cli/install\">pnpm install | pnpm</a></li>\n<li><a href=\"https://nodejs.org/api/corepack.html\">Corepack | Node.js Documentation</a></li>\n<li><a href=\"https://semver.org/lang/zh-CN/\">Semantic Versioning 2.0.0</a></li>\n</ul>\n","date_published":"2022-01-11T00:00:00.000Z","date_modified":"2026-08-01T00:00:00.000Z","tags":["前端","js","npm","pnpm","lockfile","可信构建"],"language":"zh"},{"id":"https://www.lihuanyu.com/posts/2021/21%E5%B9%B4%E9%9A%8F%E7%AC%94/","url":"https://www.lihuanyu.com/posts/2021/21%E5%B9%B4%E9%9A%8F%E7%AC%94/","title":"21年随笔","summary":"记录 2021 年回到成都后的工作、装修、新家、疫情、行业变化、技术状态、读书、换手机和生活节点。","content_html":"<blockquote>\n<p>21年快要过完了，再不写点什么，可真就什么都没写了。水一篇，记录生活。</p>\n</blockquote>\n<h2>城市</h2>\n<p>20年10月换了工作，辞别待了4年的北京，回到了成都。</p>\n<p>21年一整年，几乎都是在成都度过的。</p>\n<p>一切看起来没什么区别，又有很多说不清的不同。</p>\n<p>照样繁忙，甚至可以说更繁忙的工作，工作日基本天天两点一线，早10晚9，和在北京，别无二致。工作的地点倒是相对繁华了不少，从“穷乡僻壤”西二旗改到了高楼遍地的天府X街。（然鹅了解下两边的房价，就知道谁才是真正的穷乡僻壤了 :) ）</p>\n<p>物价/吃饭等方面，其实也和北京差不多，成都的南门，确实不太像成都。</p>\n<p>但是吃穿住行，后面三个还是有不同的。</p>\n<p>北京的房租，一个开间一个月5500，成都一个两居室一个月只要2600。</p>\n<p>北京打车，动辄30，打个百来块钱的车，非常正常。成都打车，很少能超过30。</p>\n<h2>装修</h2>\n<p>回家后安排了装修，差不多小半年的时间完成了装修，一切从简。less is more是真的，简单的才是最耐看最好打理最好维护的，就像代码，越短，bug越少，哈。</p>\n<p>此处应该有图？没认真拍过满意的，先占位吧。</p>\n<h2>新家</h2>\n<p>所以，在搞定装修后，再散味3个月左右。</p>\n<p>是的，住进自家的房子啦。开心。</p>\n<p>难道这就是所谓的归属感？可能吧。</p>\n<h2>疫情</h2>\n<p>21年一整年，新冠疫情仍然笼罩于世界。</p>\n<p>去哪都不太方便，主要就去了两次杭州出差，去了一次西双版纳团建。杭州也是个大工地，到处挖得破破烂烂的，修地铁修路。早期城市规划人员恐怕想都不敢想现在城市的规模/状态。西双版纳有点意思，景色很棒，尤其是夜景和当地的特色服装。</p>\n<p>21年初的春节，很多地方倡导/要求就地过年，回到成都在这种背景下显得很明智，跟父母近了很多。希望今年的春节，团聚的人能多一些。</p>\n<h2>行业</h2>\n<p>21年的前端技术，感觉确实没那么精彩了，都在搞一些修修补补，贴近业务的一些优化。</p>\n<p>同时能明显感觉到经济大环境在恶化，退守二线城市也是希望在这个充满不确定的世界中不要翻车。</p>\n<p>年底了，中概互联还在跌跌不休，裁员的消息一波接一波，各路互联网公司还在消减福利。也不知道未来如何发展。</p>\n<p>不过互联网还算好的，教培行业直接欢声笑语中打出GG。</p>\n<h2>技术</h2>\n<p>感觉有点懈怠，这一年无所长进。也许是回家，要处理的生活事务太多了吧。</p>\n<p>不过很有意思，服务器里托管着我的blog和mpxjs的文档，一年多的时间，几乎没有出过任何问题。</p>\n<p>我的blog可能没有更新不算什么有难度的事情，但mpxjs的文档是在持续迭代的。</p>\n<p>仅仅靠着GitHub Action与certbot的自动脚本，就能持续保证文档的更新，很有意思。</p>\n<h2>图书</h2>\n<p>我不是一个爱读书的人，之前的很多技术书籍，都没有认真看过，可能就翻了一两页吧。于是后来也不怎么买书了。</p>\n<p>回到成都后有了充足的书房空间，开始搞一些奇奇怪怪的书籍，比如《中国是部金融史》、《高效人士的秘诀》。前一本是历史书还挺有意思的，历史确实像是在螺旋中前进，前人犯过的错，后人也要跟着踩一遍坑。</p>\n<p>技术书籍买了一本《JavaScript悟道》，文笔很有意思，希望这次能读完。</p>\n<h2>手机</h2>\n<p>用了N年的小米，换成了苹果。</p>\n<p>契机主要是有两个。一是后续考虑买车用车，carplay有显著的优势。二是小米11的火龙888烧了WiFi，并且小米的解决方案令我非常不满意，有欺骗消费者的嫌疑。so，用脚投票。</p>\n<p>说实话，换过来发现在某些使用体验方面，小米真的做得超级棒。iPhone的一些设计，很奇怪，很反直觉。比如侧滑返回，苹果是做不到一直返回的。再比如一些本土化功能，也非常不贴心。</p>\n<p>但反过来，小米确实不配做高端机。半年后价格大跳水，没有核心科技。机器稳定性很差，我是一个用东西很爱惜的人，所以多年使用小米并觉得很好用，但从不敢向不太熟悉的朋友推荐小米，因为真的容易用坏。电池不耐用，系统套路多，如果你没有一点geek精神，会拿到一部广告机。</p>\n<p>苹果现在用下来比较舒服的地方，拍照尤其是拍人，很好看。电池非常耐用，终于明白之前用iPhone的朋友们动不动掏出只有20%不到的电的手机还一点不慌是怎么回事了。以及，不用曲面屏，真的是个加分项。</p>\n<h2>one more thing</h2>\n<p>领证了，以后也是有家室的人了。希望未来能经营好我们的小家庭。</p>\n<p>21年再见，22年你好。</p>\n","date_published":"2021-12-30T00:00:00.000Z","tags":["生活","随笔"],"language":"zh"},{"id":"https://www.lihuanyu.com/en/posts/2020/didi-mini-program-i18n-engineering/","url":"https://www.lihuanyu.com/en/posts/2020/didi-mini-program-i18n-engineering/","title":"Internationalization for Large Mini Programs: Engineering Lessons from Didi","summary":"A review of Didi Mini Program internationalization, covering copy governance, Mini Program runtime constraints, WXS-based translation, cross-platform adaptation, and team workflow.","content_html":"<p>In 2020, Didi Mini Program needed an English version. At first glance, this sounded like translating Chinese strings into English. In practice, it was a full engineering and collaboration project.</p>\n<p>The Mini Program had many business lines, shared libraries, frontend hardcoded copy, and server-delivered text. The launch date was fixed, frontend staffing was limited, and translation, integration, testing, and release all had to happen in one coordinated flow.</p>\n<p>The English version launched on schedule and ran stably. The point of this review is not to prove that one framework feature is powerful. It is to summarize what large Mini Programs really need when they add internationalization.</p>\n<p><a href=\"/posts/2020/%E6%BB%B4%E6%BB%B4%E5%87%BA%E8%A1%8C%E5%B0%8F%E7%A8%8B%E5%BA%8FI18n%E6%9C%80%E4%BD%B3%E5%AE%9E%E8%B7%B5/\">Chinese version of this article</a></p>\n<p><img src=\"https://dpubstatic.udache.com/static/dpubimg/F3Ev6M08X7/clipboard_image_1599113935176.png\" alt=\"Didi Mini Program i18n\"></p>\n<h2>Internationalization Is Content and Runtime Governance</h2>\n<p>i18n is short for internationalization. For an application, it is not only translation. It includes:</p>\n<ul>\n<li>How copy is collected and named.</li>\n<li>How translation resources are maintained.</li>\n<li>How templates and JavaScript read text consistently.</li>\n<li>How dates, times, numbers, and currency are formatted.</li>\n<li>How UI updates when the locale changes.</li>\n<li>How server-delivered copy works with frontend copy.</li>\n<li>How multiple business lines follow the same convention.</li>\n</ul>\n<p>For a large Mini Program, the difficulty is not one <code>$t('hello')</code> call. The difficulty is governance at scale: many business lines, pages, platforms, and teams. Any temporary convention becomes expensive later.</p>\n<h2>Treat Copy as an Asset First</h2>\n<p>The first step should not be code changes. It should be copy inventory.</p>\n<p>Copy usually comes from several sources:</p>\n<ol>\n<li>Static text in frontend templates.</li>\n<li>Toasts, dialogs, and error messages in JavaScript.</li>\n<li>Default copy inside component libraries and shared libraries.</li>\n<li>Server-delivered campaign, order, status, and operation text.</li>\n<li>Text embedded in images, icons, empty states, and marketing assets.</li>\n</ol>\n<p>Without inventory, the team will keep finding pages where most text is English but one dialog remains Chinese.</p>\n<p>A reusable approach:</p>\n<ul>\n<li>Give every text item a stable key instead of using Chinese text as the key.</li>\n<li>Organize keys by domain and page, such as <code>order.detail.cancelTitle</code>.</li>\n<li>Decide which copy belongs to frontend language packs and which is delivered by the server according to locale.</li>\n<li>Review new copy during code review to prevent new hardcoded strings.</li>\n<li>Define fallback behavior so missing keys do not produce blank UI.</li>\n</ul>\n<p>Once copy has structure, translation, testing, and incremental maintenance become manageable.</p>\n<h2>Mini Program Runtime Constraints Matter</h2>\n<p>In Web applications, calling JavaScript functions from template expressions is natural. Mini Programs are different.</p>\n<p>Taking WeChat Mini Program as an example, the runtime separates the logic layer and rendering layer:</p>\n<ul>\n<li>JavaScript runs in the logic layer.</li>\n<li>The rendering layer displays the page.</li>\n<li>Data is passed from logic to rendering through <code>setData</code>.</li>\n<li>The rendering layer cannot execute normal JavaScript.</li>\n<li>WXS is a view-layer scripting capability.</li>\n</ul>\n<p>This creates an important issue: if every translation happens in the logic layer and then gets passed to the template through <code>setData</code>, locale changes and list rendering can increase cross-thread communication.</p>\n<p>An i18n solution has to respect runtime boundaries, not only API design.</p>\n<h2>Why Template Translation Functions Matter</h2>\n<p>The ideal usage should feel close to Web i18n:</p>\n<pre><code class=\"language-html\">&lt;template&gt;\n  &lt;view&gt;{{ $t('message.hello', { name: userName }) }}&lt;/view&gt;\n  &lt;view&gt;{{ formattedDatetime }}&lt;/view&gt;\n&lt;/template&gt;\n</code></pre>\n<p>JavaScript should use the same capability:</p>\n<pre><code class=\"language-js\">import mpx, { createComponent } from '@mpxjs/core'\n\ncreateComponent({\n  ready () {\n    console.log(this.$t('message.hello', { name: 'Didi' }))\n    this.$i18n.locale = 'en-US'\n  },\n  computed: {\n    formattedDatetime () {\n      return this.$d(new Date(), 'long')\n    }\n  }\n})\n</code></pre>\n<p>This API looks simple, but two problems sit behind it:</p>\n<ol>\n<li>Can the template execute a translation function directly?</li>\n<li>Can JavaScript reuse the same language pack and formatting logic?</li>\n</ol>\n<p>Mpx solves this by generating WXS translation functions at build time from the language dictionaries, then injecting them into templates that use translation calls. On the JavaScript side, the corresponding logic is transformed and injected into the runtime.</p>\n<p>Templates and JavaScript can then share a unified i18n API while avoiding unnecessary cross-thread data transfer.</p>\n<h2>Language Packs Belong in the Build System</h2>\n<p>A language pack configuration can look like this:</p>\n<pre><code class=\"language-js\">new MpxWebpackPlugin({\n  i18n: {\n    locale: 'en-US',\n    messages: {\n      'en-US': {\n        message: {\n          hello: '{name} world'\n        }\n      },\n      'zh-CN': {\n        message: {\n          hello: '{name} 世界'\n        }\n      }\n    }\n  }\n})\n</code></pre>\n<p>Language packs can be inline, but large projects should keep them as separate modules because translation resources need review, testing, and continuous maintenance.</p>\n<p>Putting language packs into the build system has several benefits:</p>\n<ul>\n<li>Templates, JavaScript, and components use the same resources.</li>\n<li>Build steps can check whether keys exist.</li>\n<li>Language pack changes do not bypass the application build.</li>\n<li>Multiple platform outputs can share one configuration.</li>\n<li>Date, number, plural, and other formatting features can be extended later.</li>\n</ul>\n<p>If i18n resources live outside the build flow, teams can easily update a language pack but forget to regenerate some intermediate artifact.</p>\n<h2>Cross-Platform Differences Should Stay in the Framework</h2>\n<p>Mini Program internationalization has another challenge: view-layer scripting differs across platforms.</p>\n<p>WeChat has WXS. Alipay has SJS. Other platforms have their own syntax and runtime restrictions. Business developers should not handle these differences in every page.</p>\n<p>Mpx absorbs this at the framework and build-system level. It can use WeChat WXS as a DSL, parse and transform it during build, then output scripts that different platforms can understand. Template-side and JavaScript-side i18n both build on this capability.</p>\n<p>The reusable principle is: <strong>cross-platform differences should be absorbed by the framework and build system, not leaked into business code.</strong></p>\n<p>The closer business code stays to one unified API, the cheaper future locale and platform expansion becomes.</p>\n<h2>Tradeoffs Compared with Other Approaches</h2>\n<p>One approach is to compute translated text in the logic layer and pass it into templates. This is easy to understand, but it increases <code>setData</code> communication. In list rendering, it can also enlarge data transfer. It may work for small projects, but it becomes expensive in large and complex pages.</p>\n<p>Another approach is the official WeChat i18n solution, which also uses view-layer scripting. But if the surrounding build process, JavaScript injection, cross-platform adaptation, and reactive locale updates are not unified, integration can still feel fragmented.</p>\n<p>The value of Mpx is not only that it can translate text. It connects the whole flow:</p>\n<ul>\n<li>Language packs are injected at build time.</li>\n<li>Templates can call translation functions directly.</li>\n<li>JavaScript uses the same API.</li>\n<li>Locale changes can trigger reactive updates.</li>\n<li>Cross-platform output is handled by the framework.</li>\n<li>Web output can reuse an experience close to vue-i18n.</li>\n</ul>\n<p>For large projects, the key question is whether the whole chain is closed, not whether one API exists.</p>\n<h2>Reusable Methodology</h2>\n<p>The Didi i18n work can be summarized into several steps.</p>\n<h3>1. Inventory Copy and Define Ownership</h3>\n<p>Classify all text by source: frontend static copy, server-delivered dynamic copy, component-library copy, and marketing asset copy. Decide ownership before changing code.</p>\n<h3>2. Design Stable Keys</h3>\n<p>Chinese source text changes. Business copy changes. Keys should describe domain meaning and location, not depend on the current wording.</p>\n<h3>3. Put Language Resources into Build and Review</h3>\n<p>Every new page, component, toast, and dialog should add language keys. Code review should catch new hardcoded copy.</p>\n<h3>4. Choose Translation Execution Location Based on Runtime</h3>\n<p>Web, Mini Program, React Native, and Flutter have different runtimes. Where translation executes affects performance and maintainability. Mini Programs especially need to consider communication between logic and rendering layers.</p>\n<h3>5. Keep One Business API</h3>\n<p>Business developers should use capabilities such as <code>$t</code>, <code>$d</code>, and <code>$n</code>. They should not need to care about WXS, SJS, or platform-specific runtime details.</p>\n<h3>6. Turn Testing into a Product Checklist</h3>\n<p>i18n testing should cover:</p>\n<ul>\n<li>First screen and core workflows.</li>\n<li>Toasts, dialogs, error states, and empty states.</li>\n<li>Long English text causing wrapping or truncation.</li>\n<li>Date, time, amount, and units.</li>\n<li>Server-delivered copy.</li>\n<li>UI updates after locale switching.</li>\n<li>Differences across platform outputs.</li>\n</ul>\n<h3>7. Accept That i18n Affects Product Design</h3>\n<p>Chinese text is short. English can be much longer. Some languages have plural forms. Some regions need different phrasing. Internationalization is not a final translation layer. It pushes component layout, copy length, and information structure to become more resilient.</p>\n<h2>Conclusion</h2>\n<p>The core lesson from Didi Mini Program’s English version was not “choose an i18n library.” It was to treat internationalization as an engineering system.</p>\n<p>Copy needs asset management. Language packs need to enter the build. Templates and JavaScript need one API. Cross-platform differences need to be absorbed by the framework. Testing needs to cover real business paths.</p>\n<p>For large Mini Programs, the hard part of i18n is not the translation function itself. It is scaled collaboration and runtime constraints. Once those two are handled, multilingual support stops being a long-term maintenance burden.</p>\n","date_published":"2020-08-31T00:00:00.000Z","date_modified":"2026-05-04T00:00:00.000Z","tags":["Mini Program","i18n","Internationalization","Mpx","Engineering"],"language":"en"},{"id":"https://www.lihuanyu.com/posts/2020/%E6%BB%B4%E6%BB%B4%E5%87%BA%E8%A1%8C%E5%B0%8F%E7%A8%8B%E5%BA%8FI18n%E6%9C%80%E4%BD%B3%E5%AE%9E%E8%B7%B5/","url":"https://www.lihuanyu.com/posts/2020/%E6%BB%B4%E6%BB%B4%E5%87%BA%E8%A1%8C%E5%B0%8F%E7%A8%8B%E5%BA%8FI18n%E6%9C%80%E4%BD%B3%E5%AE%9E%E8%B7%B5/","title":"大型小程序国际化实践：滴滴出行 i18n 的工程方法","summary":"复盘滴滴出行小程序英文版改造，从文案治理、双线程架构、WXS 翻译函数、跨平台适配和协作流程中提炼大型小程序国际化方法论。","content_html":"<p>2020 年，滴滴出行小程序需要支持英文版。这个需求看起来是“把中文换成英文”，真正落地时却是一个完整的工程协作问题。</p>\n<p>当时小程序里有大量业务线、公共库、前端硬编码文案和服务端下发文案。英文版上线时间明确，前端投入有限，翻译、联调、测试、发布都要在同一条链路里完成。</p>\n<p>最后英文版按期上线并稳定运行。这篇文章复盘的重点，不是证明某个框架功能有多强，而是总结大型小程序做国际化时真正需要处理的几类问题。</p>\n<p><a href=\"/en/posts/2020/didi-mini-program-i18n-engineering/\">English version: Internationalization for Large Mini Programs</a></p>\n<p><img src=\"https://dpubstatic.udache.com/static/dpubimg/F3Ev6M08X7/clipboard_image_1599113935176.png\" alt=\"滴滴出行微信小程序 i18n\"></p>\n<h2>国际化不是翻译，是内容和运行时治理</h2>\n<p>i18n 是 Internationalization 的缩写，指软件具备支持多语言、多地区格式的能力。对应用来说，它不只是把中文翻成英文，还包括：</p>\n<ul>\n<li>文案如何收集和命名。</li>\n<li>翻译资源如何维护。</li>\n<li>模板和 JS 中如何统一取文案。</li>\n<li>日期、时间、数字、货币如何格式化。</li>\n<li>语言切换后界面如何响应更新。</li>\n<li>服务端下发文案如何和前端文案协同。</li>\n<li>多业务线如何在同一套规范下接入。</li>\n</ul>\n<p>大型小程序的国际化难点，不在单个 <code>$t('hello')</code> 调用，而在规模化治理：业务线多、页面多、平台多、团队多，任何临时约定都会在后续维护中放大成本。</p>\n<h2>先把文案当资产管理</h2>\n<p>国际化项目的第一步，不应该是改代码，而是清点文案。</p>\n<p>文案通常来自几类地方：</p>\n<ol>\n<li>前端模板里的静态文本。</li>\n<li>JS 逻辑里的 toast、弹窗、错误提示。</li>\n<li>组件库和公共库里的默认文案。</li>\n<li>服务端下发的活动、订单、状态和运营文案。</li>\n<li>图片、图标、空状态和营销素材里的文字。</li>\n</ol>\n<p>如果不先清点，后面就会出现大量“页面已经切英文，但某个弹窗还是中文”的问题。</p>\n<p>可复用的做法是：</p>\n<ul>\n<li>给每条文案稳定 key，而不是直接用中文做 key。</li>\n<li>key 按业务域和页面层级组织，例如 <code>order.detail.cancelTitle</code>。</li>\n<li>明确哪些文案归前端语言包，哪些由服务端按 locale 下发。</li>\n<li>把新增文案纳入代码审查，避免继续写硬编码。</li>\n<li>对兜底语言做明确约定，防止 key 缺失时页面空白。</li>\n</ul>\n<p>文案一旦有了结构，翻译、测试和后续增量维护才会变成可管理流程。</p>\n<h2>小程序运行时带来的特殊挑战</h2>\n<p>Web 应用里，模板表达式调用 JS 函数很自然。但小程序不是标准浏览器运行时。</p>\n<p>以微信小程序为例，它采用逻辑层和渲染层分离的架构：</p>\n<ul>\n<li>逻辑层运行 JavaScript。</li>\n<li>渲染层负责页面展示。</li>\n<li>数据通过 <code>setData</code> 从逻辑层传到渲染层。</li>\n<li>渲染层不能直接执行普通 JS。</li>\n<li>WXS 是运行在视图层的类 JS 脚本能力。</li>\n</ul>\n<p>这带来一个关键问题：如果所有翻译都在逻辑层完成，再通过 <code>setData</code> 把结果传给模板，语言切换或列表渲染时就会增加线程通信成本。</p>\n<p>国际化方案必须考虑运行时边界，而不是只考虑 API 是否好看。</p>\n<h2>为什么模板翻译函数很重要</h2>\n<p>理想使用方式应该接近 Web 里的 i18n 体验：</p>\n<pre><code class=\"language-html\">&lt;template&gt;\n  &lt;view&gt;{{ $t('message.hello', { name: userName }) }}&lt;/view&gt;\n  &lt;view&gt;{{ formattedDatetime }}&lt;/view&gt;\n&lt;/template&gt;\n</code></pre>\n<p>在 JS 中也应该能使用同一套能力：</p>\n<pre><code class=\"language-js\">import mpx, { createComponent } from '@mpxjs/core'\n\ncreateComponent({\n  ready () {\n    console.log(this.$t('message.hello', { name: 'Didi' }))\n    this.$i18n.locale = 'en-US'\n  },\n  computed: {\n    formattedDatetime () {\n      return this.$d(new Date(), 'long')\n    }\n  }\n})\n</code></pre>\n<p>这个 API 看起来简单，背后要解决两个问题：</p>\n<ol>\n<li>模板里能不能直接执行翻译函数。</li>\n<li>JS 里能不能复用同一份语言包和格式化逻辑。</li>\n</ol>\n<p>Mpx 的做法，是在构建阶段把语言字典和翻译函数合成可在视图层执行的 WXS，并自动注入到使用翻译函数的模板中。JS 侧则通过框架能力把对应翻译逻辑转换并注入到逻辑层运行时。</p>\n<p>这样模板和 JS 都可以使用统一的 i18n API，同时减少不必要的跨线程数据传递。</p>\n<h2>语言包应该在构建体系里统一管理</h2>\n<p>语言包配置通常长这样：</p>\n<pre><code class=\"language-js\">new MpxWebpackPlugin({\n  i18n: {\n    locale: 'en-US',\n    messages: {\n      'en-US': {\n        message: {\n          hello: '{name} world'\n        }\n      },\n      'zh-CN': {\n        message: {\n          hello: '{name} 世界'\n        }\n      }\n    }\n  }\n})\n</code></pre>\n<p>语言包既可以直接写在配置里，也可以独立成模块路径。大型项目更适合后者，因为语言资源需要被翻译、审查、测试和持续维护。</p>\n<p>把语言包纳入统一构建体系有几个好处：</p>\n<ul>\n<li>模板、JS、组件都使用同一份资源。</li>\n<li>构建时可以检查 key 是否存在。</li>\n<li>语言包更新不会脱离应用构建流程。</li>\n<li>多端产物可以共享同一套配置。</li>\n<li>后续可以扩展日期、数字、复数等格式化能力。</li>\n</ul>\n<p>国际化一旦脱离构建体系，就容易变成“改了语言包但忘记重新生成某份中间产物”的维护问题。</p>\n<h2>跨平台适配要藏在框架层</h2>\n<p>小程序国际化还有一个特殊难点：不同平台的视图层脚本能力并不完全一样。</p>\n<p>微信有 WXS，支付宝有 SJS，百度、QQ、字节等平台也有自己的语法和运行限制。业务开发者不应该在每个页面里手写一套平台差异处理。</p>\n<p>Mpx 的跨平台能力在这里发挥了作用：以微信 WXS 作为 DSL，在构建阶段解析、转换，再输出到不同平台可识别的脚本形式。模板侧和 JS 侧的 i18n 能力都建立在这套转换能力上。</p>\n<p>这背后的方法论是：<strong>跨平台差异应该被框架和构建系统吸收，而不是泄露给业务代码。</strong></p>\n<p>业务代码越接近统一 API，后续语言扩展和平台扩展的成本越低。</p>\n<h2>和其他方案的取舍</h2>\n<p>当时也对比过其他思路。</p>\n<p>一种方案是利用 computed，把翻译结果在逻辑层算好，再传给模板。这个方案理解成本低，但会增加 <code>setData</code> 通信，列表场景里还可能放大数据传输量。它适合小项目，但在大型复杂页面里会让性能和维护成本变高。</p>\n<p>另一种方案是微信官方的 i18n 方案，思路也使用视图层脚本。但如果周边构建、JS 注入、跨平台适配和响应式能力没有统一起来，业务接入仍然需要处理更多零散环节。</p>\n<p>Mpx 的优势不只是“能翻译”，而是把几个环节串成一条链：</p>\n<ul>\n<li>构建时注入语言包。</li>\n<li>模板中直接使用翻译函数。</li>\n<li>JS 中使用同一套 API。</li>\n<li>locale 变化可以响应式更新。</li>\n<li>跨平台输出由框架统一处理。</li>\n<li>Web 产物可以复用类似 vue-i18n 的体验。</li>\n</ul>\n<p>大型项目选方案时，应该优先看整条链路是否闭合，而不是只比较单个 API。</p>\n<h2>可复用的方法论</h2>\n<p>大型小程序 i18n 改造可以抽象成几个步骤。</p>\n<h3>1. 先做文案盘点和归属划分</h3>\n<p>把所有文案按来源分类：前端静态文案、服务端动态文案、组件库文案、运营素材文案。先定归属，再改代码。</p>\n<h3>2. 设计稳定 key，而不是依赖中文原文</h3>\n<p>中文原文会修改，业务文案会调整。key 应该表达业务含义和位置，不能直接依赖当前文案。</p>\n<h3>3. 把语言资源纳入构建和审查</h3>\n<p>新增页面、新增组件、新增 toast，都应该同步新增语言 key。代码审查时要看是否还有硬编码文案。</p>\n<h3>4. 根据运行时选择翻译执行位置</h3>\n<p>Web、小程序、React Native、Flutter 的运行时不同，翻译函数放在哪里执行会影响性能和维护成本。小程序尤其要考虑逻辑层和渲染层通信。</p>\n<h3>5. 让业务代码使用统一 API</h3>\n<p>业务开发者应该只关心 <code>$t</code>、<code>$d</code>、<code>$n</code> 这类能力，不应该关心 WXS、SJS 或平台差异。</p>\n<h3>6. 把测试清单产品化</h3>\n<p>国际化测试不能只靠看几个页面。至少要覆盖：</p>\n<ul>\n<li>首屏和核心流程。</li>\n<li>toast、弹窗、错误态、空状态。</li>\n<li>长英文导致的换行和截断。</li>\n<li>日期、时间、金额、单位。</li>\n<li>服务端下发文案。</li>\n<li>语言切换后的页面刷新。</li>\n<li>多平台产物差异。</li>\n</ul>\n<h3>7. 接受国际化会反推产品设计</h3>\n<p>中文短，英文长；中文没有复数，英文有复数；部分文案在不同地区表达方式不同。国际化不是最后套一层翻译，它会反过来要求组件布局、文案长度和信息结构更稳健。</p>\n<h2>总结</h2>\n<p>滴滴出行小程序英文版改造的核心经验，不是“找一个 i18n 库”，而是把国际化当成工程系统来做。</p>\n<p>文案要有资产管理，语言包要进入构建，模板和 JS 要使用统一 API，跨平台差异要被框架吸收，测试要覆盖真实业务路径。</p>\n<p>对大型小程序来说，国际化的难点从来不是翻译函数本身，而是规模化协作和运行时约束。只要把这两件事处理好，多语言支持就不会变成后续迭代里的负担。</p>\n","date_published":"2020-08-31T00:00:00.000Z","date_modified":"2026-05-04T00:00:00.000Z","tags":["小程序","i18n","国际化","Mpx","工程化"],"language":"zh"},{"id":"https://www.lihuanyu.com/en/posts/2020/github-actions-automation-entry-point-not-deployment-machine/","url":"https://www.lihuanyu.com/en/posts/2020/github-actions-automation-entry-point-not-deployment-machine/","title":"GitHub Actions Is an Automation Entry Point, Not a Deployment Machine","summary":"A practical review of where GitHub Actions works well, where it becomes the wrong execution environment, and how to draw a cleaner boundary for small project deployments.","content_html":"<p>My first experience with CI/CD was Travis CI on GitHub open source projects. Later, Travis became less attractive for personal projects, while GitHub Actions was built directly into the repository workflow. Blogs, documentation sites, and small open source projects naturally moved there.</p>\n<p>The judgment at the time was simple: if the code lives on GitHub, automation should live next to it. Push code, run tests, build artifacts, publish packages, deploy the site. The experience was smooth.</p>\n<p>Looking back after a few years, that judgment was only half right.</p>\n<p>GitHub Actions is excellent for one-off automation around a repository: testing, linting, building, publishing npm packages, building Docker images, generating documentation, and notifying external systems. It is less suitable as the default execution environment for every deployment, especially when the target server is far away from the runner, the deployment needs to transfer many files, the process depends on server-local state, or the operational boundary should be clearer.</p>\n<p><a href=\"/posts/2020/%E4%BB%8ETravis%E8%BF%81%E7%A7%BB%E5%88%B0GitHub-Actions/\">Chinese version of this article</a></p>\n<p>This article revisits several related experiences:</p>\n<ul>\n<li>Moving from Travis CI to GitHub Actions.</li>\n<li>Publishing npm packages with Actions.</li>\n<li>Building Docker images in Actions.</li>\n<li>Changing this blog’s deployment from “Actions uploads the built files” to “Actions sends a signed webhook, and the server pulls, builds, and publishes locally.”</li>\n</ul>\n<p>The short version is: <strong>GitHub Actions is a good automation entry point, but it should not automatically become the production deployment machine.</strong></p>\n<h2>CI Belongs Close to the Repository</h2>\n<p>Travis CI was attractive in the early days because the setup was simple. Open source code was already on GitHub, and a <code>.travis.yml</code> file could run tests and builds after every push.</p>\n<p>The downside was that CI lived in another system. Permissions, logs, triggers, caching, and deployment behavior all had to be understood across two platforms. Once GitHub Actions matured, putting CI back beside the repository became the natural choice.</p>\n<p>Actions has several strengths:</p>\n<ol>\n<li>It is connected to GitHub events such as <code>push</code>, <code>pull_request</code>, <code>release</code>, and <code>workflow_dispatch</code>.</li>\n<li>Secrets, permissions, environments, and branch protection live in the same platform.</li>\n<li>The Marketplace covers many common tasks.</li>\n<li>Hosted runners cover Ubuntu, Windows, and macOS.</li>\n<li>Logs, checks, and pull request gates are part of the code review flow.</li>\n</ol>\n<p>In an Mpx template project, I used a matrix to test generated projects across operating systems and Node.js versions:</p>\n<pre><code class=\"language-yaml\">strategy:\n  matrix:\n    os: [macos-latest, windows-latest, ubuntu-latest]\n    node: [10, 12, 14]\n</code></pre>\n<p>That kind of verification is hard to do on one local machine and easy to do on cloud runners. Template projects are especially sensitive to “the generated project does not run,” and Actions can catch that kind of regression on every commit.</p>\n<p>So CI is the strongest use case for GitHub Actions: <strong>if a task is stateless, repeatable, and tightly connected to repository code, it is usually a good fit.</strong></p>\n<h2>Good Fit: Tests, Lint, and Builds</h2>\n<p>This is the least controversial use case.</p>\n<pre><code class=\"language-yaml\">name: test\n\non:\n  pull_request:\n  push:\n    branches:\n      - master\n\njobs:\n  test:\n    runs-on: ubuntu-latest\n    steps:\n      - uses: actions/checkout@v4\n      - uses: actions/setup-node@v4\n        with:\n          node-version: 24\n          cache: pnpm\n      - run: corepack enable\n      - run: pnpm install --frozen-lockfile\n      - run: pnpm run lint\n      - run: pnpm test\n      - run: pnpm run build\n</code></pre>\n<p>These jobs have a clean shape:</p>\n<ul>\n<li>The input is repository code and the lockfile.</li>\n<li>The output is a test result or build artifact.</li>\n<li>A failure can block a merge.</li>\n<li>The job does not depend on production server state.</li>\n<li>It can be rerun safely.</li>\n</ul>\n<p>Within this boundary, Actions adds clear value: quality gates become automatic instead of relying on someone remembering to run commands locally.</p>\n<h2>Good Fit: Publishing npm Packages</h2>\n<p>Publishing npm packages also fits GitHub Actions well because it is essentially a transformation from repository state to a registry version.</p>\n<p>A stable pattern is tag-based publishing:</p>\n<pre><code class=\"language-yaml\">name: publish\n\non:\n  push:\n    tags:\n      - 'v*'\n\njobs:\n  publish:\n    runs-on: ubuntu-latest\n    permissions:\n      contents: read\n      id-token: write\n    steps:\n      - uses: actions/checkout@v4\n      - uses: actions/setup-node@v4\n        with:\n          node-version: 24\n          registry-url: https://registry.npmjs.org/\n      - run: corepack enable\n      - run: pnpm install --frozen-lockfile\n      - run: pnpm test\n      - run: pnpm run build\n      - run: npm publish --provenance --access public\n</code></pre>\n<p>Two details matter here.</p>\n<p>First, publishing should include tests and builds. A publish workflow is not only <code>npm publish</code>; it should encode the definition of “publishable.”</p>\n<p>Second, npm trusted publishing should be the preferred direction when possible. It uses OIDC to establish trust between GitHub Actions and npm, reducing the need for long-lived npm tokens. If a project has not adopted trusted publishing yet, an npm automation token can still be a transitional option.</p>\n<p>This use case works because Actions can connect tags, builds, tests, versions, and publish logs in one flow. The runner is temporary, but the publishing job should be temporary too.</p>\n<h2>Good Fit: Building and Pushing Docker Images</h2>\n<p>Docker image builds are also often a good fit, especially when the image is pushed to Docker Hub, GitHub Container Registry, or another registry.</p>\n<p>A typical workflow looks like this:</p>\n<pre><code class=\"language-yaml\">name: docker\n\non:\n  push:\n    branches:\n      - main\n\njobs:\n  build:\n    runs-on: ubuntu-latest\n    steps:\n      - uses: actions/checkout@v4\n      - uses: docker/setup-buildx-action@v3\n      - uses: docker/login-action@v3\n        with:\n          username: ${{ secrets.DOCKERHUB_USERNAME }}\n          password: ${{ secrets.DOCKERHUB_TOKEN }}\n      - uses: docker/build-push-action@v6\n        with:\n          push: true\n          tags: your-name/your-image:latest\n</code></pre>\n<p>The advantages are straightforward:</p>\n<ul>\n<li>The build environment is clean.</li>\n<li>Images can be tagged with a version, tag, or commit SHA.</li>\n<li>The server only needs to pull and run the image.</li>\n<li>Developers do not all need a complete Docker build environment locally.</li>\n</ul>\n<p>The limit is runner resources. Normal web service images are usually fine. AI-related images can be different: Python, CUDA, PyTorch, and model dependencies can consume disk space quickly. I once built images related to A1111/Stable-Diffusion-WebUI and had to clean unused tools and caches from the runner before the build could finish.</p>\n<p>That kind of cleanup can work, but it is not an infinite scaling strategy. If images keep growing, better options include:</p>\n<ul>\n<li>Optimizing the Dockerfile and removing unnecessary dependencies.</li>\n<li>Using BuildKit cache or registry cache.</li>\n<li>Using GitHub larger runners.</li>\n<li>Using self-hosted runners.</li>\n<li>Moving the build closer to the target environment or to a dedicated build service.</li>\n</ul>\n<p>Docker builds fit Actions as long as the resource scale still fits the runner. Once the workflow depends on forcing enough disk space out of a hosted runner every time, the boundary is already being stretched.</p>\n<h2>Poor Fit: Uploading Large Deployments to a Remote Server</h2>\n<p>This blog’s deployment change is a useful example of where Actions can become the wrong execution environment.</p>\n<p>The earlier deployment flow was simple: GitHub Actions built the blog and synchronized the generated static files to the server. It worked for a while.</p>\n<p>The problem was distance. The runner was overseas, while the server was in China. The build itself was not slow. Most of the time was spent transferring files across an unstable path. One deployment took nearly four minutes, and a large part of that time had little to do with application logic.</p>\n<p>This kind of setup has several risks:</p>\n<ul>\n<li>Network quality between the runner and the server is not under your control.</li>\n<li>More static files mean more transfer time.</li>\n<li>SSH keys or deployment tokens have to live in GitHub Secrets.</li>\n<li>Failures require checking both Actions logs and server state.</li>\n<li>Nginx, certificates, process managers, local caches, and directory permissions are not truly controlled by Actions.</li>\n</ul>\n<p>The deployment flow now looks like this:</p>\n<pre><code class=\"language-text\">git push\n  -&gt; GitHub Actions\n  -&gt; send a signed webhook\n  -&gt; a NestJS service on the target server receives the notification\n  -&gt; the server runs git pull / pnpm install / build locally\n  -&gt; the result is switched into the Nginx site directory\n</code></pre>\n<p>Actions now does one light job: notification.</p>\n<p>The workflow core is roughly:</p>\n<pre><code class=\"language-yaml\">name: deploy\n\non:\n  push:\n    branches:\n      - master\n  workflow_dispatch:\n\nconcurrency:\n  group: production-deploy\n  cancel-in-progress: true\n\njobs:\n  notify:\n    runs-on: ubuntu-latest\n    steps:\n      - name: Notify deployment server\n        env:\n          DEPLOY_WEBHOOK_URL: ${{ secrets.DEPLOY_WEBHOOK_URL }}\n          DEPLOY_WEBHOOK_SECRET: ${{ secrets.DEPLOY_WEBHOOK_SECRET }}\n          REPOSITORY: ${{ github.repository }}\n          REF: ${{ github.ref }}\n          SHA: ${{ github.sha }}\n        run: |\n          payload=&quot;$(jq -cn \\\n            --arg repository &quot;$REPOSITORY&quot; \\\n            --arg ref &quot;$REF&quot; \\\n            --arg sha &quot;$SHA&quot; \\\n            '{ repository: $repository, ref: $ref, sha: $sha }')&quot;\n\n          signature=&quot;sha256=$(printf '%s' &quot;$payload&quot; \\\n            | openssl dgst -sha256 -hmac &quot;$DEPLOY_WEBHOOK_SECRET&quot; -binary \\\n            | xxd -p -c 256)&quot;\n\n          curl --fail-with-body --request POST &quot;$DEPLOY_WEBHOOK_URL&quot; \\\n            --header &quot;Content-Type: application/json&quot; \\\n            --header &quot;X-Lihuanyu-Signature-256: $signature&quot; \\\n            --data &quot;$payload&quot;\n</code></pre>\n<p>The important change is that Actions no longer carries the artifact. It sends a verifiable deployment request, and the actual deployment happens inside the target environment.</p>\n<p>The benefits are direct:</p>\n<ul>\n<li>The GitHub Actions job becomes shorter.</li>\n<li>Large file transfers from an overseas runner to a domestic server disappear.</li>\n<li>The server can reuse local Git, pnpm cache, and build environment.</li>\n<li>Deployment logs live closer to Nginx, PM2, certificates, and system state.</li>\n<li>GitHub Secrets only need a webhook URL and signing secret, not a server SSH private key.</li>\n<li>The webhook can verify repository, branch, commit SHA, and HMAC signature.</li>\n</ul>\n<p>There are costs too:</p>\n<ul>\n<li>The server must maintain Node.js, pnpm, Git, build scripts, and deployment directory permissions.</li>\n<li>The webhook service needs authentication, locks, logs, and error handling.</li>\n<li>The initial clone or a bad GitHub network path can still be slow from the server side.</li>\n<li>Concurrent pushes must be serialized.</li>\n<li>Rollback has to be designed in the server deployment script.</li>\n</ul>\n<p>But these are deployment-system concerns anyway. Handling them on the target server is closer to the real runtime environment.</p>\n<h2>Poor Fit: Long-Lived Operational State</h2>\n<p>GitHub-hosted runners are temporary machines for workflow jobs. They are not a place to keep important state.</p>\n<p>That makes them a poor fit for tasks such as:</p>\n<ul>\n<li>Saving important data outside normal build caches.</li>\n<li>Running operations that depend on local machine state.</li>\n<li>Performing tasks that need long manual observation.</li>\n<li>Putting database migrations, service restarts, certificate updates, and directory switching into one fragile remote script.</li>\n</ul>\n<p>Database migrations, Nginx reloads, PM2 reloads, certificate renewals, and static directory switches can be triggered by CI/CD. But the execution logic should usually live in the target environment, with clear logs, locks, and failure handling.</p>\n<p>Actions can start a deployment. It does not always need to execute the deployment.</p>\n<h2>Poor Fit: Giving Secrets to Untrusted Code</h2>\n<p>Publishing, deployment, and image pushing all involve secrets.</p>\n<p>The common risk is mixing untrusted code with powerful secrets. Workflows triggered by forked pull requests, third-party actions without pinned versions, dynamically downloaded scripts, and broad <code>GITHUB_TOKEN</code> permissions can all expand the blast radius.</p>\n<p>My default rules are:</p>\n<ul>\n<li>Set explicit minimum <code>permissions</code>.</li>\n<li>Run publishing jobs only on tags, releases, or protected branches.</li>\n<li>Run deployment jobs only from the main branch or manual triggers.</li>\n<li>Pin third-party actions to clear versions; for critical paths, consider pinning to commit SHAs.</li>\n<li>Prefer npm trusted publishing over long-lived npm tokens.</li>\n<li>Prefer signed deployment webhooks over handing a server SSH key to every workflow.</li>\n</ul>\n<p>Actions is good at automation, and automation means repeating something reliably. If the permission boundary is wrong, it also repeats the mistake reliably.</p>\n<h2>A Simple Decision Table</h2>\n<div class=\"table-scroll\"><table>\n<thead>\n<tr>\n<th>Scenario</th>\n<th>Fit for GitHub Actions</th>\n<th>Judgment</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Lint, unit tests, type checks</td>\n<td>Good</td>\n<td>Stateless, repeatable, directly useful for code review.</td>\n</tr>\n<tr>\n<td>Matrix tests across OS and Node versions</td>\n<td>Good</td>\n<td>Hosted runners cover platforms that are hard to reproduce locally.</td>\n</tr>\n<tr>\n<td>Static site builds</td>\n<td>Good</td>\n<td>Artifacts are clear and failures are cheap.</td>\n</tr>\n<tr>\n<td>npm package publishing</td>\n<td>Good</td>\n<td>Tags, tests, builds, and publishing can form one closed loop.</td>\n</tr>\n<tr>\n<td>Docker image build and push</td>\n<td>Usually good</td>\n<td>Very large images may need larger runners, self-hosted runners, or dedicated builders.</td>\n</tr>\n<tr>\n<td>GitHub Pages / Cloudflare Pages deployment</td>\n<td>Good</td>\n<td>The target platform is close to Actions and the flow is standardized.</td>\n</tr>\n<tr>\n<td>Uploading many files to a distant server</td>\n<td>Not ideal</td>\n<td>Network path and transfer time can dominate the deployment.</td>\n</tr>\n<tr>\n<td>Production local builds and directory switching</td>\n<td>Better on the server</td>\n<td>The execution is closer to the real environment, with clearer logs and permissions.</td>\n</tr>\n<tr>\n<td>Database migrations and service restarts</td>\n<td>Be careful</td>\n<td>Actions can trigger them, but execution needs locks, logs, and rollback behavior.</td>\n</tr>\n<tr>\n<td>Tasks requiring fixed IP or private network access</td>\n<td>Depends</td>\n<td>Larger runners, self-hosted runners, or server-side execution may be better.</td>\n</tr>\n</tbody>\n</table>\n</div><h2>How I Design Small Project Deployment Now</h2>\n<p>For a personal blog, admin tool, or small service, I would split responsibilities this way.</p>\n<p>GitHub Actions handles:</p>\n<ol>\n<li>Tests, type checks, and builds during pull requests.</li>\n<li>Deployment notification after a push to the main branch.</li>\n<li>Standard artifact publishing, such as npm packages, Docker images, or documentation.</li>\n<li>Logging the commit SHA, actor, and workflow run URL.</li>\n</ol>\n<p>The server handles:</p>\n<ol>\n<li>Verifying webhook signature, repository, branch, and commit SHA.</li>\n<li>Serializing deployments to avoid overlapping writes.</li>\n<li>Pulling code, installing dependencies, and running the build.</li>\n<li>Switching artifacts into the Nginx site directory.</li>\n<li>Recording deployment logs and supporting rollback when necessary.</li>\n<li>Managing Node.js, pnpm, PM2, Nginx, certificates, and system permissions.</li>\n</ol>\n<p>This division is not complicated, but the boundary is cleaner: GitHub Actions acts as the trigger and quality gate, while the server executes production deployment inside the production-like environment.</p>\n<h2>Conclusion</h2>\n<p>When I first moved from Travis CI to GitHub Actions, I mostly cared about stability and convenience. That judgment still holds: GitHub Actions is a very good automation entry point for personal projects and open source projects.</p>\n<p>After using it for npm publishing, Docker image building, and blog deployment, the boundary is clearer.</p>\n<p>Tasks that are strongly tied to repository state, stateless, repeatable, and easy to rerun belong in Actions: tests, builds, package publishing, image building, documentation generation, and deployment notifications.</p>\n<p>Tasks that depend on production server state, require large file transfers, involve long-lived operational permissions, or need detailed runtime handling should not automatically be pushed into Actions. For those cases, Actions is often better as the trigger than as the execution environment.</p>\n<p>In one sentence: <strong>treat GitHub Actions as an automation entry point, not as the only deployment machine.</strong></p>\n<h2>Further Reading</h2>\n<ul>\n<li><a href=\"https://docs.github.com/actions/reference/specifications-for-github-hosted-runners\">GitHub Docs: GitHub-hosted runners</a></li>\n<li><a href=\"https://docs.github.com/actions/concepts/workflows-and-actions/concurrency\">GitHub Docs: Concurrency</a></li>\n<li><a href=\"https://docs.github.com/actions/tutorials/publish-packages/publish-nodejs-packages\">GitHub Docs: Publishing Node.js packages</a></li>\n<li><a href=\"https://docs.npmjs.com/trusted-publishers\">npm Docs: Trusted publishing for npm packages</a></li>\n<li><a href=\"https://docs.github.com/en/actions/concepts/runners/larger-runners\">GitHub Docs: Larger runners</a></li>\n</ul>\n","date_published":"2020-06-21T00:00:00.000Z","date_modified":"2026-05-04T00:00:00.000Z","tags":["GitHub Actions","CI CD","Deployment","Automation","Webhook"],"language":"en"},{"id":"https://www.lihuanyu.com/posts/2020/%E4%BB%8ETravis%E8%BF%81%E7%A7%BB%E5%88%B0GitHub-Actions/","url":"https://www.lihuanyu.com/posts/2020/%E4%BB%8ETravis%E8%BF%81%E7%A7%BB%E5%88%B0GitHub-Actions/","title":"GitHub Actions 适合做什么，不适合做什么","summary":"从 Travis 迁移、npm 自动发布、Docker 镜像构建和博客 webhook 部署改造出发，重新梳理 GitHub Actions 的使用边界。","content_html":"<p>我最早用 CI/CD，是在 GitHub 开源项目里接入 Travis。后来 Travis 稳定性和免费策略都不再适合个人项目，GitHub Actions 又和 GitHub 仓库天然集成，于是博客、文档和开源项目陆续迁了过去。</p>\n<p>那时我的判断很简单：代码在 GitHub，自动化也放在 GitHub，提交后自动测试、构建、发布，体验非常顺。</p>\n<p>几年后再看，这个判断只对了一半。</p>\n<p>GitHub Actions 很适合做“围绕仓库的一次性自动化任务”：测试、构建、发布 npm 包、构建 Docker 镜像、生成文档、通知外部系统。它不一定适合直接完成所有部署，尤其是目标服务器在国内、需要传输大量静态文件、依赖服务器本地状态或需要更清晰运维边界时。</p>\n<p>这篇文章把几段实践放在一起复盘：</p>\n<ul>\n<li>从 Travis 迁移到 GitHub Actions。</li>\n<li>用 Actions 自动发布 npm 包。</li>\n<li>在 Actions 里构建大型 Docker 镜像。</li>\n<li>把博客部署从“Actions 直接部署”改成“Actions 通知服务器，服务器本地拉取、构建、发布”。</li>\n</ul>\n<p>结论先放前面：<strong>GitHub Actions 是很好的自动化入口，但不应该默认成为生产服务器的执行环境。</strong></p>\n<p><a href=\"/en/posts/2020/github-actions-automation-entry-point-not-deployment-machine/\">English version: GitHub Actions Is an Automation Entry Point, Not a Deployment Machine</a></p>\n<h2>从 Travis 到 GitHub Actions：CI 最适合放在仓库旁边</h2>\n<p>早期用 Travis 的原因很简单：开源项目托管在 GitHub，Travis 接入方便，写一个 <code>.travis.yml</code> 就能在每次提交后跑测试和构建。</p>\n<p>但 Travis 最大的问题是稳定性和生态集成。CI 结果在另一个系统里，权限、日志、触发条件、缓存和部署都要跨系统理解。后来 GitHub Actions 成熟后，把 CI 放回 GitHub 仓库附近就很自然。</p>\n<p>Actions 的优势主要有几类：</p>\n<ol>\n<li>触发条件和 GitHub 事件天然打通，比如 <code>push</code>、<code>pull_request</code>、<code>release</code>、<code>workflow_dispatch</code>。</li>\n<li>Secrets、权限、环境和分支保护都在 GitHub 里管理。</li>\n<li>Marketplace 里有大量可复用 action，常见任务不用从零写脚本。</li>\n<li>标准 runner 覆盖 Ubuntu、Windows、macOS，适合做跨平台验证。</li>\n<li>日志、状态检查、PR 门禁都和代码审查流程在一起。</li>\n</ol>\n<p>以前在 Mpx 模板项目里，我就用过 matrix 同时覆盖多个操作系统和 Node 版本：</p>\n<pre><code class=\"language-yaml\">strategy:\n  matrix:\n    os: [macos-latest, windows-latest, ubuntu-latest]\n    node: [10, 12, 14]\n</code></pre>\n<p>这种事情放在本机很难做，放在云端 runner 非常合适。模板项目最怕的是“生成出来的项目不能跑”，而 Actions 可以在每次提交后把不同平台、不同 Node 版本的基础构建都跑一遍。</p>\n<p>所以，CI 是 GitHub Actions 最稳的基本盘：<strong>只要任务是无状态的、可重复的、和仓库代码强相关的，就很适合放在 Actions 里。</strong></p>\n<h2>适合场景一：测试、Lint 和构建</h2>\n<p>这是最没有争议的场景。</p>\n<pre><code class=\"language-yaml\">name: test\n\non:\n  pull_request:\n  push:\n    branches:\n      - master\n\njobs:\n  test:\n    runs-on: ubuntu-latest\n    steps:\n      - uses: actions/checkout@v4\n      - uses: actions/setup-node@v4\n        with:\n          node-version: 24\n          cache: pnpm\n      - run: corepack enable\n      - run: pnpm install --frozen-lockfile\n      - run: pnpm run lint\n      - run: pnpm test\n      - run: pnpm run build\n</code></pre>\n<p>这类任务有几个共同点：</p>\n<ul>\n<li>输入是仓库代码和 lockfile。</li>\n<li>输出是测试结果或构建产物。</li>\n<li>失败可以阻止合并。</li>\n<li>不依赖生产服务器状态。</li>\n<li>可以安全地重复执行。</li>\n</ul>\n<p>在这个边界里，Actions 的价值很明确：让质量门禁自动化，不靠人记得执行命令。</p>\n<h2>适合场景二：发布 npm 包</h2>\n<p>npm 包发布也适合放在 Actions 里，因为它本质上是“仓库状态 -&gt; registry 版本”的自动化。</p>\n<p>比较稳的触发方式是 tag 发布：</p>\n<pre><code class=\"language-yaml\">name: publish\n\non:\n  push:\n    tags:\n      - 'v*'\n\njobs:\n  publish:\n    runs-on: ubuntu-latest\n    permissions:\n      contents: read\n      id-token: write\n    steps:\n      - uses: actions/checkout@v4\n      - uses: actions/setup-node@v4\n        with:\n          node-version: 24\n          registry-url: https://registry.npmjs.org/\n      - run: corepack enable\n      - run: pnpm install --frozen-lockfile\n      - run: pnpm test\n      - run: pnpm run build\n      - run: npm publish --provenance --access public\n</code></pre>\n<p>这里有两个关键点。</p>\n<p>第一，发布前必须跑测试和构建。发包不是单纯执行 <code>npm publish</code>，而是把“可发布”的判断写进流程。</p>\n<p>第二，今天更应该优先考虑 npm trusted publishing，也就是用 OIDC 建立 GitHub Actions 和 npm 之间的信任关系，减少长期 npm token 的暴露面。如果项目暂时还没配置 trusted publishing，再用 npm automation token 作为过渡方案。</p>\n<p>发布 npm 包适合放在 Actions 里，是因为 Actions 能把版本、tag、构建、测试、发布日志连在一起。它的执行环境虽然是临时的，但发布任务本来也应该是临时的。</p>\n<h2>适合场景三：构建并推送 Docker 镜像</h2>\n<p>Docker 镜像构建也经常适合放在 Actions 里，尤其是镜像需要推到 Docker Hub、GitHub Container Registry 或其他镜像仓库时。</p>\n<p>典型流程是：</p>\n<pre><code class=\"language-yaml\">name: docker\n\non:\n  push:\n    branches:\n      - main\n\njobs:\n  build:\n    runs-on: ubuntu-latest\n    steps:\n      - uses: actions/checkout@v4\n      - uses: docker/setup-buildx-action@v3\n      - uses: docker/login-action@v3\n        with:\n          username: ${{ secrets.DOCKERHUB_USERNAME }}\n          password: ${{ secrets.DOCKERHUB_TOKEN }}\n      - uses: docker/build-push-action@v6\n        with:\n          push: true\n          tags: your-name/your-image:latest\n</code></pre>\n<p>这个场景里，Actions 的好处也很明显：</p>\n<ul>\n<li>构建环境干净。</li>\n<li>可以结合 tag 或 commit sha 标记镜像。</li>\n<li>可以把镜像推到 registry，让服务器只负责拉镜像和运行。</li>\n<li>不需要每个开发者本机都有完整 Docker 构建环境。</li>\n</ul>\n<p>但 Docker 镜像构建会碰到 runner 资源边界。普通 Web 服务镜像通常没问题；如果是 Stable Diffusion 这类 AI 镜像，Python、CUDA、PyTorch 和模型相关依赖会迅速吃掉磁盘空间。以前我构建过 A1111/Stable-Diffusion-WebUI 相关镜像，需要先清理 runner 上不需要的工具和缓存，才能勉强构建成功。</p>\n<p>这种技巧有用，但不是无限扩展方案。镜像继续变大后，更合理的选择是：</p>\n<ul>\n<li>优化 Dockerfile，减少层和无用依赖。</li>\n<li>使用 BuildKit cache 或 registry cache。</li>\n<li>使用 GitHub larger runners。</li>\n<li>使用 self-hosted runner。</li>\n<li>把构建放到更靠近目标环境的机器或专用构建服务里。</li>\n</ul>\n<p>所以 Docker 构建适合 Actions，但前提是资源规模仍在 runner 能承受的范围内。超过这个范围，就不应该继续靠清理磁盘硬撑。</p>\n<h2>不适合场景一：把国内服务器部署完全压在 Actions 上</h2>\n<p>博客部署改造，就是 GitHub Actions 使用边界的一个典型案例。</p>\n<p>之前的部署思路是：GitHub Actions 负责构建博客，然后把生成的静态文件同步到服务器。这个方案简单直观，早期也能工作。</p>\n<p>问题是，Actions 的 runner 在海外，目标服务器在国内。构建本身不慢，慢的是把文件从 runner 传到服务器。一次部署接近 4 分钟，其中很多时间并没有花在业务逻辑上，而是花在网络传输和远程同步上。</p>\n<p>这类场景有几个隐患：</p>\n<ul>\n<li>runner 到服务器的网络不可控。</li>\n<li>静态文件越多，传输越容易成为瓶颈。</li>\n<li>SSH key 或部署 token 要放在 GitHub Secrets 里。</li>\n<li>部署失败时，需要同时查 Actions 日志和服务器状态。</li>\n<li>服务器本地环境、Nginx、证书、进程管理都不是 Actions 真正能掌控的东西。</li>\n</ul>\n<p>所以现在博客部署改成了另一种结构：</p>\n<pre><code class=\"language-text\">git push\n  -&gt; GitHub Actions\n  -&gt; 发送带签名的 webhook\n  -&gt; 目标服务器上的 NestJS 接收通知\n  -&gt; 服务器本地 git pull / pnpm install / build\n  -&gt; 同步到 Nginx 站点目录\n</code></pre>\n<p>Actions 现在只负责一件很轻的事：通知。</p>\n<p>当前 workflow 的核心大概是这样：</p>\n<pre><code class=\"language-yaml\">name: deploy\n\non:\n  push:\n    branches:\n      - master\n  workflow_dispatch:\n\nconcurrency:\n  group: production-deploy\n  cancel-in-progress: true\n\njobs:\n  notify:\n    runs-on: ubuntu-latest\n    steps:\n      - name: Notify deployment server\n        env:\n          DEPLOY_WEBHOOK_URL: ${{ secrets.DEPLOY_WEBHOOK_URL }}\n          DEPLOY_WEBHOOK_SECRET: ${{ secrets.DEPLOY_WEBHOOK_SECRET }}\n          REPOSITORY: ${{ github.repository }}\n          REF: ${{ github.ref }}\n          SHA: ${{ github.sha }}\n        run: |\n          payload=&quot;$(jq -cn \\\n            --arg repository &quot;$REPOSITORY&quot; \\\n            --arg ref &quot;$REF&quot; \\\n            --arg sha &quot;$SHA&quot; \\\n            '{ repository: $repository, ref: $ref, sha: $sha }')&quot;\n\n          signature=&quot;sha256=$(printf '%s' &quot;$payload&quot; \\\n            | openssl dgst -sha256 -hmac &quot;$DEPLOY_WEBHOOK_SECRET&quot; -binary \\\n            | xxd -p -c 256)&quot;\n\n          curl --fail-with-body --request POST &quot;$DEPLOY_WEBHOOK_URL&quot; \\\n            --header &quot;Content-Type: application/json&quot; \\\n            --header &quot;X-Lihuanyu-Signature-256: $signature&quot; \\\n            --data &quot;$payload&quot;\n</code></pre>\n<p>这个方案的变化点在于：Actions 不再搬运产物，只发送一个可验证的部署请求。真正的部署发生在目标服务器上。</p>\n<p>这样做的收益很直接：</p>\n<ul>\n<li>GitHub Actions 执行时间变短。</li>\n<li>不再从海外 runner 向国内服务器传大量文件。</li>\n<li>服务器可以复用本地 Git、pnpm 缓存和构建环境。</li>\n<li>部署日志集中在目标服务器，和 Nginx、PM2、证书状态更接近。</li>\n<li>GitHub Secrets 里不需要保存服务器 SSH 私钥，只保存 webhook URL 和签名密钥。</li>\n<li>webhook 可以校验仓库、分支、commit sha 和 HMAC 签名，避免被随便触发。</li>\n</ul>\n<p>代价也要承认：</p>\n<ul>\n<li>服务器上要维护 Node、pnpm、Git、构建脚本和部署目录权限。</li>\n<li>webhook 服务要有鉴权、锁、日志和错误处理。</li>\n<li>首次 clone 或 GitHub 网络波动时，服务器拉代码也可能慢。</li>\n<li>部署过程需要处理并发推送，避免两个部署互相覆盖。</li>\n<li>回滚要在服务器部署脚本里设计，而不是只看 Actions。</li>\n</ul>\n<p>但这些代价属于部署系统本来就应该处理的问题。把它们放在目标服务器上，反而更接近真实运行环境。</p>\n<h2>不适合场景二：需要长期状态的运维动作</h2>\n<p>GitHub-hosted runner 是临时环境。GitHub 官方文档也把 hosted runner 描述为执行 workflow job 的机器，通常每次任务都是新环境。</p>\n<p>这意味着它不适合承担长期状态：</p>\n<ul>\n<li>不适合保存构建缓存以外的重要数据。</li>\n<li>不适合做依赖本机状态的运维任务。</li>\n<li>不适合在 runner 上做需要人工持续观察的操作。</li>\n<li>不适合把数据库迁移、服务重启、证书更新等全部交给一段远程脚本硬跑。</li>\n</ul>\n<p>数据库迁移、Nginx reload、PM2 reload、证书续期、静态目录切换，这些都可以被 CI/CD 触发，但最好由目标环境里的部署脚本负责执行，并且有清晰日志和失败处理。</p>\n<p>Actions 可以发起部署，不一定要亲自执行部署。</p>\n<h2>不适合场景三：把 Secrets 暴露给不可信代码</h2>\n<p>只要涉及发布、部署、镜像推送，就会涉及 secrets。</p>\n<p>常见风险是：workflow 既能跑不可信代码，又能拿到高权限 secrets。比如来自 fork 的 PR、动态下载的 action、没有固定版本的第三方 action、过宽的 <code>GITHUB_TOKEN</code> 权限，都可能扩大风险。</p>\n<p>我的默认策略是：</p>\n<ul>\n<li><code>permissions</code> 显式写最小权限。</li>\n<li>发布任务只在 tag、release 或受保护分支上触发。</li>\n<li>部署任务只接受主分支或手动触发。</li>\n<li>第三方 action 固定到明确版本，关键场景可进一步固定到 commit sha。</li>\n<li>npm 发布优先使用 OIDC trusted publishing，减少长期 token。</li>\n<li>服务器部署优先用 webhook 签名，不把 SSH 私钥直接交给所有 workflow。</li>\n</ul>\n<p>Actions 很适合自动化，但自动化的本质是“稳定地重复执行”。如果权限边界没想清楚，它也会稳定地重复放大错误。</p>\n<h2>一张简单判断表</h2>\n<div class=\"table-scroll\"><table>\n<thead>\n<tr>\n<th>场景</th>\n<th>是否适合 GitHub Actions</th>\n<th>判断</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Lint、单元测试、类型检查</td>\n<td>适合</td>\n<td>无状态、可重复、直接服务代码审查。</td>\n</tr>\n<tr>\n<td>多系统、多 Node 版本矩阵测试</td>\n<td>适合</td>\n<td>runner 平台丰富，本机很难覆盖。</td>\n</tr>\n<tr>\n<td>静态站点构建</td>\n<td>适合</td>\n<td>构建产物明确，失败成本低。</td>\n</tr>\n<tr>\n<td>npm 包发布</td>\n<td>适合</td>\n<td>tag、测试、构建、发布可以形成闭环。</td>\n</tr>\n<tr>\n<td>Docker 镜像构建并推送 registry</td>\n<td>通常适合</td>\n<td>镜像太大时要考虑 larger runner、自托管 runner 或专用构建环境。</td>\n</tr>\n<tr>\n<td>部署到 GitHub Pages / Cloudflare Pages</td>\n<td>适合</td>\n<td>平台离 Actions 近，流程标准化。</td>\n</tr>\n<tr>\n<td>向国内服务器传大量文件</td>\n<td>不太适合</td>\n<td>网络链路和传输耗时容易成为瓶颈。</td>\n</tr>\n<tr>\n<td>生产服务器上的本地构建与目录切换</td>\n<td>更适合放服务器</td>\n<td>更接近真实环境，日志和权限更清楚。</td>\n</tr>\n<tr>\n<td>数据库迁移和服务重启</td>\n<td>谨慎</td>\n<td>可以由 Actions 触发，但执行逻辑应有锁、回滚和日志。</td>\n</tr>\n<tr>\n<td>需要固定 IP 或内网访问的任务</td>\n<td>看情况</td>\n<td>larger runner、自托管 runner 或服务器本地执行更合适。</td>\n</tr>\n</tbody>\n</table>\n</div><h2>我现在会怎么设计个人项目部署</h2>\n<p>如果是个人博客、管理后台、小型服务，我现在会按下面的方式分工。</p>\n<p>GitHub Actions 负责：</p>\n<ol>\n<li>PR 阶段跑测试、类型检查和构建。</li>\n<li>主分支 push 后发送部署通知。</li>\n<li>npm 包、Docker 镜像、文档这类标准产物的发布。</li>\n<li>在日志里记录 commit sha、触发人、workflow run URL。</li>\n</ol>\n<p>服务器负责：</p>\n<ol>\n<li>校验 webhook 签名、仓库、分支和 commit sha。</li>\n<li>串行化部署，避免并发覆盖。</li>\n<li>拉取代码，安装依赖，执行构建。</li>\n<li>把产物切换到 Nginx 站点目录。</li>\n<li>记录部署日志，必要时支持回滚。</li>\n<li>管理 Node、pnpm、PM2、Nginx、证书和系统权限。</li>\n</ol>\n<p>这个分工并不复杂，但边界更合理：GitHub Actions 做“触发器”和“质量门禁”，服务器做“生产环境里的部署执行者”。</p>\n<h2>总结</h2>\n<p>从 Travis 迁移到 GitHub Actions 时，我更关注的是 CI/CD 能不能更稳定、更方便。这个判断今天仍然成立：GitHub Actions 是个人项目和开源项目非常好用的自动化入口。</p>\n<p>但经历过 npm 发布、Docker 镜像构建和博客部署改造后，我对它的边界更清楚了。</p>\n<p>适合放在 Actions 里的，是和仓库强相关、无状态、可重复、失败后容易重跑的任务。比如测试、构建、发包、镜像构建、文档生成和部署通知。</p>\n<p>不适合完全压在 Actions 上的，是依赖生产服务器状态、需要大量跨境传输、涉及长期权限和运维细节的任务。对这些场景，Actions 更适合作为触发器，而不是执行环境本身。</p>\n<p>一句话总结：<strong>把 GitHub Actions 当自动化入口，不要把它当唯一的部署机器。</strong></p>\n<h2>扩展阅读</h2>\n<ul>\n<li><a href=\"https://docs.github.com/actions/reference/specifications-for-github-hosted-runners\">GitHub Docs: GitHub-hosted runners</a></li>\n<li><a href=\"https://docs.github.com/actions/concepts/workflows-and-actions/concurrency\">GitHub Docs: Concurrency</a></li>\n<li><a href=\"https://docs.github.com/actions/tutorials/publish-packages/publish-nodejs-packages\">GitHub Docs: Publishing Node.js packages</a></li>\n<li><a href=\"https://docs.npmjs.com/trusted-publishers\">npm Docs: Trusted publishing for npm packages</a></li>\n<li><a href=\"https://docs.github.com/en/actions/concepts/runners/larger-runners\">GitHub Docs: Larger runners</a></li>\n</ul>\n","date_published":"2020-06-21T00:00:00.000Z","date_modified":"2026-05-04T00:00:00.000Z","tags":["GitHub Actions","CI CD","持续集成","自动化部署","Webhook"],"language":"zh"},{"id":"https://www.lihuanyu.com/en/posts/2020/didi-mini-program-package-size-optimization/","url":"https://www.lihuanyu.com/en/posts/2020/didi-mini-program-package-size-optimization/","title":"Package Size Governance for Large Mini Programs: Lessons from Didi","summary":"A review of Didi Mini Program package size optimization, covering size budgets, dependency analysis, subpackages, npm dependency placement, and architecture tradeoffs.","content_html":"<p>In the second half of 2019, Didi needed to migrate the WebApp entry inside WeChat Wallet and Alipay’s grid menu into Mini Programs. This was not a simple wrapper change. Ride hailing, bus, designated driving, bike, hitch, car services, and other business lines all had to live inside one Mini Program.</p>\n<p>The first major engineering problem was package size.</p>\n<p>Mini Program platforms impose package size limits, especially on the main package and each subpackage. Didi’s home page also carried many high-frequency workflows: choosing a service, entering origin and destination, switching car types, keeping state, and entering orders. The more business logic concentrated on the home page, the more code was pulled into the main package.</p>\n<p><a href=\"/posts/2020/%E6%BB%B4%E6%BB%B4%E5%87%BA%E8%A1%8C%E5%B0%8F%E7%A8%8B%E5%BA%8F%E4%BD%93%E7%A7%AF%E4%BC%98%E5%8C%96%E5%AE%9E%E8%B7%B5/\">Chinese version of this article</a></p>\n<p>This article is not only a list of optimizations. It is a review of a reusable method: when a Mini Program grows from one business line into a multi-business, multi-team, dependency-heavy application, package size control has to become governance rather than occasional cleanup.</p>\n<h2>Define the Problem First</h2>\n<p>When package size is close to the limit, the first reaction is often to delete code, compress images, or enable minification. These are useful, but they only address the surface.</p>\n<p>Large Mini Program package size problems usually come from three sources:</p>\n<ol>\n<li><strong>Asset size</strong>: images, videos, fonts, JSON, static configuration, and other resources included in the package.</li>\n<li><strong>Dependency size</strong>: shared libraries, polyfills, protocol files, component libraries, and duplicated cross-business dependencies.</li>\n<li><strong>Architecture size</strong>: product structure forces many business lines into the home page, so the code cannot be delayed even if it is technically modular.</li>\n</ol>\n<p>The first two can often be improved by tooling. The third requires product, architecture, and build-system decisions. Without this distinction, teams can spend a lot of time on local optimizations without solving the main package pressure.</p>\n<h2>Step 1: Make Size Visible</h2>\n<p>Before optimization, three questions need clear answers:</p>\n<ul>\n<li>What is inside the main package?</li>\n<li>Which modules are largest?</li>\n<li>Which dependencies are duplicated or emitted to the wrong package?</li>\n</ul>\n<p>Didi Mini Program was built with Mpx, whose build pipeline is based on webpack. That made it possible to use tools such as <code>webpack-bundle-analyzer</code> to inspect the output.</p>\n<p><img src=\"https://dpubstatic.udache.com/static/dpubimg/3e488293-959e-4617-8187-69fdb532e9ab.jpg\" alt=\"package size analysis\"></p>\n<p>The value of this step is not only finding large files. It creates a shared language for multiple teams. If size discussions depend on intuition, it is hard to coordinate. When a chart shows duplicated dependencies or a subpackage-only module in the main package, the conversation becomes much more concrete.</p>\n<p>The reusable rule is simple: <strong>produce a size report before making optimization decisions.</strong></p>\n<h2>Step 2: Do Basic Optimization, But Do Not Stop There</h2>\n<p>Basic optimizations include:</p>\n<ul>\n<li>Minify JavaScript, CSS, templates, and JSON.</li>\n<li>Remove unused code and assets.</li>\n<li>Move images and videos to a CDN when possible.</li>\n<li>Use tree shaking, module deduplication, and on-demand imports.</li>\n<li>Control polyfills and shared runtime size.</li>\n<li>Avoid installing the same dependency multiple times because of version drift.</li>\n</ul>\n<p>Mpx can reuse many webpack ecosystem optimizations, and it also includes Mini Program-specific work such as page-level dependency collection, runtime compression, shared style reuse, and subpackage module extraction.</p>\n<p>These optimizations are necessary, but they usually only make the package thinner. If the product architecture forces all business logic into the main package, minification alone will not provide enough room for long-term growth.</p>\n<p>So the goal of basic optimization is to remove obvious waste and buy time for structural work.</p>\n<h2>Step 3: Move Low-Frequency Pages into Subpackages</h2>\n<p>The idea of subpackages is straightforward: pages not needed at startup should not occupy main package space. They can be downloaded when the user navigates to them.</p>\n<p>In Didi Mini Program, trip history, origin/destination selection, profile pages, and other non-home pages were early candidates for subpackages.</p>\n<p>The initial subpackage work released several hundred KB from the main package. The number was not huge, but it proved an important point: if project structure can cooperate with subpackage rules, main package size becomes manageable.</p>\n<p>Subpackage design should follow user paths:</p>\n<ul>\n<li>Startup-critical content stays in the main package.</li>\n<li>Pages reached after the first screen go into subpackages.</li>\n<li>Independent business pages go into business subpackages.</li>\n<li>Shared capabilities enter the main package only when truly shared.</li>\n<li>Modules used by only one subpackage should be emitted with that subpackage.</li>\n</ul>\n<h2>Step 4: Fix npm Dependencies That Leak into the Main Package</h2>\n<p>The difficult part is that real projects do not place all code under page directories. Many features are integrated as npm packages.</p>\n<p>Early subpackage rules often depended on file paths: files under a subpackage directory went into that subpackage, and everything else went into the main package. This worked for page files, but not for modules under <code>node_modules</code>.</p>\n<p>For example, a trip-history subpackage might use a socket library only inside that subpackage. If the library came from npm, its path was under <code>node_modules</code>. A path-only rule could still emit it into the main package.</p>\n<p>That creates a frustrating situation: business code has moved into a subpackage, but its dependencies remain in the main package.</p>\n<p>Mpx later added more precise dependency ownership analysis:</p>\n<ol>\n<li>Track which subpackages reference each module during build.</li>\n<li>Emit a module to a subpackage if only that subpackage uses it.</li>\n<li>Avoid forcing resources into the main package if they are shared only by subpackages.</li>\n<li>Generate subpackage-specific cache groups for modules reused inside the same subpackage.</li>\n</ol>\n<p>The core idea is: <strong>module ownership should be determined by usage, not only by file location.</strong></p>\n<p>This is critical for large Mini Programs. In multi-team projects, business features are often delivered as npm packages. If the build system cannot understand actual usage scope, the main package will keep absorbing dependencies that do not belong there.</p>\n<h2>Step 5: Know When Technical Optimization Reaches Its Limit</h2>\n<p>As the business kept growing, Didi Mini Program hit a harder problem: every business line needed expression on the home page.</p>\n<p>This is different from many e-commerce or content Mini Programs. An e-commerce home page can be mostly an entry point, while details, orders, search, and profile pages can be separated. A mobility home page has to carry service selection, origin and destination, car types, maps, prices, status, and recommendations. Users also expect smooth switching between services.</p>\n<p>That means each business line needs a home-page component. If the component must appear on the home page, it is hard to move it into a normal subpackage.</p>\n<p>At that stage, the main package roughly consisted of:</p>\n<ul>\n<li>Shared base libraries: framework runtime, component library, polyfills, communication libraries, and shared business dependencies.</li>\n<li>Home business code: home-page components and state logic from different business lines.</li>\n</ul>\n<p>Continuing to remove a few KB was no longer enough. The real conflict was between product architecture and package limits.</p>\n<p>That is the limit of pure technical optimization. After that point, package size work becomes an architecture decision.</p>\n<h2>Step 6: Use a Cover Page to Change Main Package Responsibility</h2>\n<p>The final solution was to make the startup page a lightweight cover page.</p>\n<p>The cover page only handled startup, brand display, and navigation. The real business home page moved into a subpackage. When users opened the Mini Program, they first entered the lightweight main package page, then navigated into the business home subpackage.</p>\n<p><img src=\"https://dpubstatic.udache.com/static/dpubimg/0f0ed782-8f67-4601-bc83-a8a343e1050b.png\" alt=\"cover page architecture\"></p>\n<p>This did not reduce total code size. It changed where code lived:</p>\n<ul>\n<li>The main package kept only startup-required and truly shared capabilities.</li>\n<li>Complex home business logic moved into a home subpackage.</li>\n<li>Future business growth mostly consumed home subpackage space instead of main package space.</li>\n</ul>\n<p>The tradeoff was clear: first business-screen display could become slower because another subpackage had to load. But compared with blocking business iteration because of main package limits, the tradeoff was acceptable. Mini Program subpackage caching also helped reduce the real user impact.</p>\n<h2>Reusable Methodology</h2>\n<p>The Didi package size work can be generalized into a sequence.</p>\n<h3>1. Set Budgets Before Hitting the Limit</h3>\n<p>Do not wait until the main package is close to the platform limit. Define budgets early:</p>\n<ul>\n<li>Main package budget.</li>\n<li>Per-subpackage budget.</li>\n<li>Shared base library budget.</li>\n<li>Per-business integration budget.</li>\n<li>Asset budgets for images, JSON, protocol files, and static resources.</li>\n</ul>\n<p>Budgets are not meant to block business. They make shared cost visible.</p>\n<h3>2. Make Every Build Show Size Changes</h3>\n<p>Package size should be monitored automatically:</p>\n<ul>\n<li>Main package and subpackage sizes.</li>\n<li>Size diff compared with the previous build.</li>\n<li>New large dependencies.</li>\n<li>Duplicated dependencies.</li>\n<li>Subpackage-only dependencies entering the main package.</li>\n</ul>\n<p>Without data, size governance becomes a one-time campaign.</p>\n<h3>3. Treat Subpackages as Architecture, Not Configuration</h3>\n<p>Subpackages are not just fields in configuration. They affect module boundaries, directory structure, npm package design, and page navigation.</p>\n<p>When a business team integrates a feature, it should answer:</p>\n<ul>\n<li>Which code is startup-critical?</li>\n<li>Which pages can be downloaded later?</li>\n<li>Will this dependency pollute the main package?</li>\n<li>Is this component truly shared?</li>\n<li>Are there unnecessary dependencies between subpackages?</li>\n</ul>\n<h3>4. Determine Dependency Ownership by Usage</h3>\n<p>In real projects, file path is not module ownership. npm packages, shared components, and utilities need to be assigned based on the dependency graph.</p>\n<p>If a module is used by only one subpackage, it should not enter the main package just because it lives under <code>node_modules</code>.</p>\n<h3>5. Product Structure Can Defeat Technical Optimization</h3>\n<p>If the home page must carry every business, the main package will grow. This cannot be solved only by minification and tree shaking.</p>\n<p>At that point, redefine the responsibility of the main package. Does it need to contain the full home page? Can it be a startup shell? Can the business home page be loaded as a subpackage? Is the user experience tradeoff acceptable?</p>\n<p>Large-scale performance work often becomes architecture work in the end.</p>\n<h2>Conclusion</h2>\n<p>Didi Mini Program package size optimization was not a set of isolated tricks. It was a staged governance path:</p>\n<ol>\n<li>Visualize package composition.</li>\n<li>Remove waste through minification, deduplication, CDN usage, and cleanup.</li>\n<li>Move low-frequency pages into subpackages.</li>\n<li>Govern npm dependencies and subpackage ownership.</li>\n<li>Change main package responsibility with a cover-page architecture when technical optimization reaches its limit.</li>\n</ol>\n<p>The most reusable lesson is the order of judgment: diagnose first, then optimize; remove waste before changing structure; solve technical issues first, then make product and architecture tradeoffs.</p>\n<p>Large Mini Programs do not stay small by accident. They need budgets, tooling, build-system support, business boundaries, and continuous monitoring.</p>\n","date_published":"2020-06-07T00:00:00.000Z","date_modified":"2026-05-04T00:00:00.000Z","tags":["Mini Program","Performance","Package Size","Mpx","Engineering"],"language":"en"},{"id":"https://www.lihuanyu.com/posts/2020/%E6%BB%B4%E6%BB%B4%E5%87%BA%E8%A1%8C%E5%B0%8F%E7%A8%8B%E5%BA%8F%E4%BD%93%E7%A7%AF%E4%BC%98%E5%8C%96%E5%AE%9E%E8%B7%B5/","url":"https://www.lihuanyu.com/posts/2020/%E6%BB%B4%E6%BB%B4%E5%87%BA%E8%A1%8C%E5%B0%8F%E7%A8%8B%E5%BA%8F%E4%BD%93%E7%A7%AF%E4%BC%98%E5%8C%96%E5%AE%9E%E8%B7%B5/","title":"大型小程序体积治理：滴滴出行的分包、依赖与架构取舍","summary":"复盘滴滴出行小程序包体积优化，从资源压缩、依赖分析、分包治理到封面页方案，整理大型小程序可复用的体积治理方法论。","content_html":"<p>2019 年下半年，滴滴出行需要把微信钱包、支付宝九宫格入口中的 WebApp 迁移为小程序。这个迁移不是简单换壳，而是要把网约车、公交、代驾、车服、单车、顺风车等业务线都接入一个统一的小程序入口。</p>\n<p>业务补齐带来的第一个工程问题，就是包体积。</p>\n<p>小程序平台对包体积有明确限制，主包和单个分包都有上限。滴滴出行小程序的首页又承载了大量高频业务：用户要在首页选择业务线、填写起终点、切换车型、保持状态、进入订单。业务越集中，首页相关代码越容易被打进主包，主包很快就会逼近平台限制。</p>\n<p><a href=\"/en/posts/2020/didi-mini-program-package-size-optimization/\">English version: Package Size Governance for Large Mini Programs</a></p>\n<p>这篇文章不只记录当时做了哪些优化，更想复盘一套可复用的方法：当一个小程序从单业务扩展到多业务、多团队、多依赖时，如何把“包体积优化”从临时救火变成长期治理。</p>\n<h2>先定义问题：不是所有体积都一样</h2>\n<p>包体积超标时，第一反应通常是“删代码”“压图片”“开压缩”。这些动作有用，但它们解决的是表层问题。</p>\n<p>大型小程序的体积问题至少分三类：</p>\n<ol>\n<li><strong>资源体积</strong>：图片、视频、字体、JSON、静态配置等资源进入包内。</li>\n<li><strong>依赖体积</strong>：公共库、polyfill、协议描述文件、组件库、跨业务基础包重复进入主包。</li>\n<li><strong>架构体积</strong>：产品信息架构让大量业务都必须挂在首页，导致代码即使按需也无法拆出去。</li>\n</ol>\n<p>前两类可以靠工程工具优化，第三类需要产品、架构和构建系统一起调整。如果没有先分清是哪一类，很容易做大量局部优化，却始终救不回主包空间。</p>\n<h2>第一步：建立体积可视化</h2>\n<p>优化体积前，必须回答三个问题：</p>\n<ul>\n<li>主包里到底有什么？</li>\n<li>哪些模块最大？</li>\n<li>哪些依赖被重复打包或被放错了位置？</li>\n</ul>\n<p>滴滴出行小程序基于 Mpx 开发，Mpx 的构建体系基于 webpack，因此可以借助 <code>webpack-bundle-analyzer</code> 一类工具分析构建产物。</p>\n<p>一个典型体积分析图会展示第三方库、公共模块和业务代码的占比：</p>\n<p><img src=\"https://dpubstatic.udache.com/static/dpubimg/3e488293-959e-4617-8187-69fdb532e9ab.jpg\" alt=\"体积分析图\"></p>\n<p>这一步的价值不只是找“大文件”，更重要的是建立团队沟通语言。体积问题如果只能靠感觉讨论，很难推动多业务线配合；一旦分析图展示出某个依赖被重复打包，或者某个只在分包使用的模块进入了主包，沟通成本会低很多。</p>\n<p>可复用的方法是：<strong>每次体积治理都先产出可视化报告，再基于报告做决策。</strong></p>\n<h2>第二步：做基础优化，但不要停在基础优化</h2>\n<p>基础优化包括：</p>\n<ul>\n<li>压缩 JS、CSS、模板和 JSON。</li>\n<li>删除无用代码和无用资源。</li>\n<li>图片、视频等静态资源尽量走 CDN。</li>\n<li>使用 tree shaking、模块去重、按需引入。</li>\n<li>控制 polyfill 和公共基础库体积。</li>\n<li>避免同一依赖因为版本不一致被打成多份。</li>\n</ul>\n<p>Mpx 基于 webpack 构建，天然能复用很多 Web 生态里的优化能力。同时，Mpx 也针对小程序做了额外处理：按页面和组件依赖收集、运行时压缩、公共样式复用、分包公共模块抽取等。</p>\n<p>样式编译同样需要明确边界。页面布局、自研组件和第三方组件不适合共享一条无差别的单位转换规则，具体选择可以参考 <a href=\"/posts/2025/%E5%B0%8F%E7%A8%8B%E5%BA%8F%E7%BB%84%E4%BB%B6%E5%BA%93%E8%AF%A5%E7%94%A8rpx%E8%BF%98%E6%98%AFpx/\">小程序组件库用 px 还是 rpx？兼容性与选型指南</a>。</p>\n<p>这些优化是必要的，但它们通常只能让包体积“变瘦”。如果业务架构决定了所有业务都要进主包，再怎么瘦身也会越来越接近上限。</p>\n<p>所以基础优化的目标不是一次性解决所有问题，而是先把明显浪费清掉，为后续架构拆分争取空间。</p>\n<h2>第三步：用分包把低频页面移出主包</h2>\n<p>小程序分包的思路很直接：启动时不需要的页面，不应该占用主包空间。用户进入对应页面时，再下载对应分包。</p>\n<p>在滴滴出行小程序里，早期比较适合拆出去的是行程页、起终点选择、个人中心等非首页页面。这些页面不是启动第一屏必须展示的内容，放到分包里对首包压力更小。</p>\n<p>初期分包完成后，主包释放了几百 KB 空间。这个收益看似不夸张，但它证明了一件事：项目结构只要能配合分包规则，主包体积就可以被持续管理。</p>\n<p>分包治理的关键不是“能拆就拆”，而是按访问路径拆：</p>\n<ul>\n<li>启动必需内容留在主包。</li>\n<li>首屏后才能访问的页面进入分包。</li>\n<li>业务线独立页面进入业务分包。</li>\n<li>公共能力只在确实公共时进入主包。</li>\n<li>只被某个分包使用的模块应该跟随分包输出。</li>\n</ul>\n<h2>第四步：治理 npm 依赖进入主包的问题</h2>\n<p>分包的难点在于，真实项目里很多代码不是按页面目录写的，而是通过 npm 包接入。</p>\n<p>早期分包规则往往依赖文件路径：分包目录下的资源进分包，其他资源进主包。这个规则对页面代码有效，但对 <code>node_modules</code> 里的业务包不友好。</p>\n<p>例如，一个行程页分包只在行程页里用到某个 socket 库，但这个库来自 npm，路径在 <code>node_modules</code> 下。如果构建系统只按路径判断，它就可能被收进主包。</p>\n<p>这会造成一个很反直觉的问题：业务代码已经拆到分包了，但业务依赖仍然留在主包。</p>\n<p>Mpx 后来做了更细的依赖归属分析：</p>\n<ol>\n<li>构建时记录每个模块被哪些分包引用。</li>\n<li>只被一个分包引用的模块输出到对应分包。</li>\n<li>被多个分包复用但不被主包使用的资源，不强行进入主包。</li>\n<li>为分包生成独立 cache group，把同一分包内复用的模块抽到分包公共 bundle。</li>\n</ol>\n<p>这类能力的核心思想是：<strong>模块归属不应该只看文件在哪，还要看它被谁使用。</strong></p>\n<p>对大型小程序来说，这一步非常关键。多团队协作时，业务常常通过 npm 包独立交付。如果构建系统不能识别 npm 依赖的真实使用范围，主包会不断吸收本不该属于它的依赖。</p>\n<h2>第五步：识别纯技术优化的边界</h2>\n<p>当业务继续增长后，滴滴出行小程序又遇到了更大的问题：所有业务线都要在首页表达需求。</p>\n<p>这和很多电商或内容类小程序不同。电商首页可以只是入口，商品详情、订单、搜索、个人中心都能拆成相对独立页面。出行首页则要同时承载业务选择、起终点、车型、地图、价格、状态、推荐等内容，用户还需要在多个业务之间流畅切换。</p>\n<p>这意味着，各业务线都要提供首页组件。只要组件必须出现在首页，它就很难被拆进普通分包。</p>\n<p>当时主包里的体积大致可以分成两块：</p>\n<ul>\n<li>公共基础库：框架运行时、组件库、polyfill、通信库、业务公共依赖。</li>\n<li>首页业务代码：各业务线在首页的需求表达组件和状态逻辑。</li>\n</ul>\n<p>这时继续做“删几 KB 代码”的收益已经不够了。真正的问题变成：产品架构要求所有业务都进入首页，而平台限制要求主包不能太大。</p>\n<p>这就是纯技术优化的边界。体积治理做到这里，必须开始讨论架构和产品形态。</p>\n<h2>第六步：用封面页方案改变主包职责</h2>\n<p>最终的解决方案，是把启动页变成一个很轻的封面页。</p>\n<p>封面页只承担启动、品牌展示和跳转职责。真正承载复杂业务的首页，被放到一个分包里。用户打开小程序后，先进入主包里的封面页，再跳转到业务首页分包。</p>\n<p><img src=\"https://dpubstatic.udache.com/static/dpubimg/0f0ed782-8f67-4601-bc83-a8a343e1050b.png\" alt=\"封面方案结构图\"></p>\n<p>这个方案没有减少总代码量，它改变的是代码位置：</p>\n<ul>\n<li>主包只保留启动必需能力和公共基础能力。</li>\n<li>复杂首页业务进入首页分包。</li>\n<li>后续业务增长主要消耗首页分包空间，而不是继续挤压主包。</li>\n</ul>\n<p>当时这个改造把一大块首页业务逻辑从主包移到了分包里，主包压力立刻缓解。更重要的是，后续业务迭代的增长位置变得可控：主包不再随着每个业务需求反复逼近上限。</p>\n<p>这个方案也有代价：首屏业务展示会变慢，因为启动后还要加载业务分包。但相比“主包超限导致无法继续上线”，这是可以接受的取舍。小程序平台本身也有分包缓存能力，实际体验可以通过加载策略继续优化。</p>\n<h2>可复用的方法论</h2>\n<p>复盘这类体积治理，大型小程序可以按下面的顺序处理类似问题。</p>\n<h3>1. 先建立预算，而不是等超限</h3>\n<p>不要等主包接近平台限制才开始治理。项目一开始就应该定义体积预算：</p>\n<ul>\n<li>主包预算。</li>\n<li>单个分包预算。</li>\n<li>公共基础库预算。</li>\n<li>单业务接入预算。</li>\n<li>图片、JSON、协议文件等资源预算。</li>\n</ul>\n<p>预算不是为了限制业务，而是为了让每个团队知道自己的代码会消耗公共空间。</p>\n<h3>2. 每次构建都能看到体积变化</h3>\n<p>体积问题适合自动化监控。至少应该能看到：</p>\n<ul>\n<li>主包和各分包大小。</li>\n<li>本次提交相比上次变化了多少。</li>\n<li>新增了哪些大依赖。</li>\n<li>是否出现重复依赖。</li>\n<li>是否有只在分包使用的依赖进入主包。</li>\n</ul>\n<p>没有数据，体积治理很容易变成临时运动。</p>\n<h3>3. 把分包当架构设计，不只是配置项</h3>\n<p>分包不只是 <code>app.json</code> 或构建配置里的一个字段。它会反过来影响业务模块边界、目录结构、npm 包设计和页面路径。</p>\n<p>大型项目里，业务团队接入时就应该回答：</p>\n<ul>\n<li>哪些代码必须首屏可用？</li>\n<li>哪些页面可以延迟下载？</li>\n<li>业务依赖是否会污染主包？</li>\n<li>公共组件是否真的公共？</li>\n<li>分包之间是否存在不合理耦合？</li>\n</ul>\n<h3>4. 依赖归属要按使用关系判断</h3>\n<p>真实项目里，文件路径并不等于模块归属。尤其是 npm 包、共享组件和公共工具函数，必须结合依赖图判断它们应该输出到哪里。</p>\n<p>一个模块如果只被某个分包使用，它就不应该因为位于 <code>node_modules</code> 而进入主包。</p>\n<h3>5. 技术优化解决不了产品结构问题</h3>\n<p>当首页必须承载所有业务时，主包天然会膨胀。这个问题不能只靠压缩和 tree shaking 解决。</p>\n<p>这时需要重新定义主包职责：主包是否真的要承载完整首页？能不能只做启动壳？业务首页是否可以作为分包加载？用户体验损失是否可接受？</p>\n<p>大型项目的性能优化，很多时候最后都会变成架构取舍。</p>\n<h2>总结</h2>\n<p>滴滴出行小程序的包体积优化，不是一组孤立技巧，而是一条逐步升级的治理路径：</p>\n<ol>\n<li>先用可视化工具看清体积组成。</li>\n<li>再做压缩、去重、CDN 化、无用代码清理。</li>\n<li>然后通过分包拆出低频页面。</li>\n<li>接着治理 npm 依赖和分包归属。</li>\n<li>最后在技术优化触顶后，用封面页方案调整主包职责。</li>\n</ol>\n<p>这套经验最值得复用的地方，不是某个具体配置，而是判断顺序：先定位，再治理；先清理浪费，再调整结构；先优化技术边界内的问题，再推动产品和架构取舍。</p>\n<p>大型小程序的包体积不会自动变好。它需要预算、工具、构建系统、业务边界和持续监控一起发挥作用。</p>\n","date_published":"2020-06-07T00:00:00.000Z","date_modified":"2026-08-01T00:00:00.000Z","tags":["小程序","性能优化","包体积","Mpx","工程化"],"language":"zh"},{"id":"https://www.lihuanyu.com/posts/2020/nginx-express%E5%81%9A%E4%B8%80%E4%B8%AA%E7%AE%80%E6%98%93%E4%BB%A3%E7%90%86%E6%9C%8D%E5%8A%A1/","url":"https://www.lihuanyu.com/posts/2020/nginx-express%E5%81%9A%E4%B8%80%E4%B8%AA%E7%AE%80%E6%98%93%E4%BB%A3%E7%90%86%E6%9C%8D%E5%8A%A1/","title":"nginx+express做一个简易代理服务","summary":"记录用 Nginx、Express、PM2 和 Let's Encrypt 搭建一个 HTTPS 代理服务的过程，以及排查 GitHub API 代理超时的细节。","content_html":"<blockquote>\n<p>用 NGINX + EXPRESS + Let’s Encrypt 构建HTTPS接口服务</p>\n</blockquote>\n<p>最近想搞个小程序，需要查询下github的api服务。小程序的请求是走的小程序提供的api，没有跨域问题，不过需要配置下请求域名白名单。</p>\n<p>本以为可以在小程序后台配置一下域名之后直接请求，结果发现小程序后台配置里只允许设备案过的域名，GitHub显然不可能有国内备案。</p>\n<p>还好自己也有服务器+域名，自己动手搭一个。而且小程序要求全是HTTPS的，还是有遇到一点小坑。</p>\n<h2>方案</h2>\n<p>其实NGINX本身就可以完成反向代理功能，不过考虑到以后可能还想在这个基础上做一些别的东西，还是来个可编程的服务比较好。出于简便性考虑直接无脑选了express。</p>\n<p>开个新的二级域名，因为只有一台云虚拟机，80/443端口已经被NGINX占用了，那就NGINX转发给express吧，所以整个结构就是NGINX反向代理express（这个express服务又是个对github的反向代理，真巧……）</p>\n<p>NGINX加一个新配置，监听80端口，对 / 全转发去本机的3000端口，完事。</p>\n<pre><code class=\"language-editorconfig\">server {\n     server_name  域名马赛克;\n     listen 80;\n \n     location / {\n         proxy_pass    http://127.0.0.1:3000;\n     }\n}\n</code></pre>\n<p>然后用Let’s Encrypt搞一下HTTPS，用 <a href=\"https://certbot.eff.org/\">certbot</a> 这个工具，直接选服务器软件和系统，可以自动帮你完成配置HTTPS。</p>\n<p>用express写个hello world，再全局安装pm2，启动我们的express程序，访问域名即可看到hello world，基本搞定。接下来再选个反向代理中间件，<a href=\"http://xn--targetapi-zb6nlqt989bjt0b.github.com\">配置下target为api.github.com</a>，大功告成。</p>\n<h2>超时</h2>\n<p>结果接下来实际测试发现炸了……</p>\n<p>很多请求报504 网关超时</p>\n<p>一通搜索，发现很多是给NGINX和PHP结合的方案用的（后仰：什么叫全宇宙最好的语言啊），不过差不多，都是说设置上游超时时间。</p>\n<p>创建 <code>/etc/nginx/conf.d/timeout.conf</code> ，加上以下内容：</p>\n<pre><code class=\"language-editorconfig\">proxy_connect_timeout       600;\nproxy_send_timeout          600;\nproxy_read_timeout          600;\nsend_timeout                600;\n</code></pre>\n<p>然后再试，发现还是容易出现504。</p>\n<p>就需要细节排查问题了，先查nginx的日志，access.log和error.log，发现这个504在nginx的错误日志里是没有的，access里有记录express程序给返回了200和504，问题应该是出在express里。</p>\n<p>接下来打印pm2的日志，发现错误信息是 <code>Error occurred while trying to proxy request xxxxx from 127.0.0.1:xxxx to https://api.github.com (ECONNREFUSED) (https://nodejs.org/api/errors.html#errors_common_system_errors)</code></p>\n<p>打开这个链接，这个错误信息是：ECONNREFUSED (Connection refused): No connection could be made because the target machine actively refused it. This usually results from trying to connect to a service that is inactive on the foreign host.</p>\n<p>目标机器主动拒绝了请求，这说明Github应该是觉得这个请求不正常，想了下估计是没带一个浏览器的头，伪装成浏览器应该就好。</p>\n<p>配置下代理的headers，再重启服务，终于OK了</p>\n","date_published":"2020-05-31T00:00:00.000Z","tags":["nodejs"],"language":"zh"},{"id":"https://www.lihuanyu.com/posts/2020/NPM%E5%B0%8F%E6%8A%80%E5%B7%A7/","url":"https://www.lihuanyu.com/posts/2020/NPM%E5%B0%8F%E6%8A%80%E5%B7%A7/","title":"npm ci 与稳定依赖安装","summary":"已并入《前端依赖、lockfile 与可信构建》。","content_html":"<p><code>npm ci</code> 的核心价值，是在已有 lockfile 的前提下，为项目做一次干净、严格、不会改写依赖描述的安装。它适合 CI、部署、排查问题、切换分支后重装依赖等场景。</p>\n<p>更完整的依赖管理实践见：</p>\n<p><a href=\"/posts/2022/%E5%89%8D%E7%AB%AF%E4%BE%9D%E8%B5%96%E4%B8%8E%E4%BF%A1%E4%BB%BB/\">前端依赖、lockfile 与可信构建</a></p>\n<p>日常开发里，<code>npm install</code> 仍然用于新增、删除、升级依赖；<code>npm ci</code> 则用于复现已经提交到仓库的依赖树。把这两个命令分清楚，能减少多人协作和云端构建里的“本地能跑、构建机不一致”问题。</p>\n","date_published":"2020-05-10T00:00:00.000Z","date_modified":"2026-05-04T00:00:00.000Z","tags":["前端","npm","lockfile"],"language":"zh"},{"id":"https://www.lihuanyu.com/posts/2019/mpx1/","url":"https://www.lihuanyu.com/posts/2019/mpx1/","title":"小程序开发者，为什么你应该尝试下MPX","summary":"从原生兼容、第三方组件支持、按需构建、跨平台编译和性能优化等角度，介绍 MPX 小程序框架的主要优势。","content_html":"<blockquote>\n<p>MPX框架 ( <a href=\"https://github.com/didi/mpx\">https://github.com/didi/mpx</a> ) 是滴滴出行推出的一款专注小程序开发的增强型框架。本篇文章将从使用角度谈谈MPX的优势与好处。如果嫌内容太长，优势部分每个小节都有简单的一句话总结，可以快速阅读。如果想了解更多设计细节，可以阅读 <a href=\"https://didi.github.io/mpx/articles/2.0.html\">前一篇文章 - MPX2.0发布</a>。</p>\n</blockquote>\n<h2>背景</h2>\n<p>在小程序逐渐火热的今天，越来越多的开发者需要进行小程序的开发。原生小程序的开发有诸多不便，开发者又需要在众多的小程序框架中做出抉择。</p>\n<p>那么今天，我们要给大家安利一款小程序框架：MPX</p>\n<h2>优势</h2>\n<p>之所以建议开发者们考虑使用MPX框架来开发小程序，是因为MPX框架具有一些别的框架所没有的优点。</p>\n<p>MPX立足原生小程序，在保证坑少的同时做了很多能力增强，提供了数据响应、模板增强、性能优化、跨平台开发等能力，以提升用户的开发体验及效率。</p>\n<p>接下来会从 原生兼容 -&gt; 第三方组件支持 -&gt; 按需构建 -&gt; 跨平台编译 -&gt; 能力增强 -&gt; 独特性能优势 六个点来逐一讲述。</p>\n<h3>原生兼容</h3>\n<blockquote>\n<p>MPX完全兼容原生，坑少。渐进接入简单。</p>\n</blockquote>\n<p>从语法风格上，我们可以看到目前市面上流行的小程序框架基本是基于web框架（taro/nanachi - react，uniapp/megalo/mpvue - vue）或者是一套全新（chameleon）/ 半全新（wepy）的标准。</p>\n<p>使用了这些框架，你所写的代码，并不是小程序代码。而是react/vue或者另一套代码。而这些代码源码到小程序代码，需要经过一次全面的转换，这个转换可能会引入一些未知的问题，产生一些坑。</p>\n<p>同时随着时间，小程序自身会逐步迭代，做出更多的功能特性，提供更好的组件、方法。而一些框架可能会受限于精力或框架节奏，没有办法第一时间跟进，甚至框架慢慢疏于维护而无法使用。</p>\n<p>而MPX选择的是，<strong>全面拥抱原生</strong>。</p>\n<p>口说无凭，我们来看个典型的MPX组件长什么样。</p>\n<p><img src=\"https://dpubstatic.udache.com/static/dpubimg/3877c653-d180-4fe7-8661-457b9de1bd82.png\" alt=\"mpx组件示例\"></p>\n<p>乍一看好像和vue没什么区别，也就多了个json块，template里写的是小程序的标签。</p>\n<p>由于这一块全是符合微信小程序原生语法，我们是不会做任何转换的，所以你写什么就是什么。（如果使用了MPX的增强特性，还是会进行一些必要的转换的，后续我们也会出文章详细解释MPX的增强是如何实现的，相对来说，我们的转换比较轻量、透明、易理解）</p>\n<p>当微信出了新的能力、新的标签、新的生命周期钩子，使用MPX框架来编写的小程序只需要直接用起来就行。</p>\n<p>所以，使用MPX框架，你可以轻易地使用 <strong><a href=\"https://developers.weixin.qq.com/miniprogram/dev/framework/custom-component/relations.html\">自定义组件的relations</a></strong> 来搞定组件间关系，使用 <strong><a href=\"https://developers.weixin.qq.com/miniprogram/dev/framework/view/wxs/\">wxs</a></strong> 来更好地构建页面。</p>\n<p>MPX几乎支持原生的每一个特性，在 .mpx 文件里，模板部分写的是原生小程序的模板语法，脚本部分写的是原生小程序的脚本语法，json部分写的是原生小程序的配置信息。用MPX，你才是真的在开发小程序。</p>\n<p>目前很多原生小程序开发者可能想尝试下框架，老项目接入框架，选MPX肯定是最简单的了。口说无凭，我们搞了个demo来给大家打个样：在我们GitHub项目中有examples文件夹，里面的 <a href=\"https://github.com/didi/mpx/tree/master/examples/mpx-progressive\">原生项目渐进接入MPX示例</a> 。</p>\n<h3>第三方组件库</h3>\n<blockquote>\n<p>MPX提供了完备的第三方组件库支持</p>\n</blockquote>\n<p>上面说了MPX对原生的极致兼容，能让你想到什么？对，就是对第三方组件库的完美支持。</p>\n<p>支持第三方组件库的重要性大家都知道，所以这个能力大部分框架都支持了。但是支持和完美支持还是有区别的。据简单观察，taro/mpvue/uniapp对于第三方组件库的支持都是以复制的形式进行的，也就是和微信小程序本来的行为很像。</p>\n<p>那么MPX是怎么支持第三方组件库的呢，这里有个demo：也在我们的GitHub里的examples文件夹下，<a href=\"https://github.com/didi/mpx/tree/master/examples/mpx-useuilib\">MPX使用第三方组件库示例</a> ，核心代码见下图：</p>\n<p><img src=\"https://dpubstatic.udache.com/static/dpubimg/9ecdd898-f7e5-48cc-89c2-bfc7ce3a03f9.png\" alt=\"MPX使用第三方组件库代码示例\"></p>\n<p>乍一眼看不太出来有什么特别的？没用过别的框架引第三方组件，简单找了找其他框架好像也没提供相应的demo，用过的朋友可以自行对比下。</p>\n<p>在MPX里使用第三方组件库，仅需要<em><strong>像web项目一样npm安装即可，并不需要复制文件</strong></em>。然后在json里直接写包名就会去node_modules下面查找了。再配合webpack alias可以做到更简单、更语义化。</p>\n<p>然而这还没有结束~</p>\n<p>细心的朋友会发现，这段示例代码中既有vant的组件，也有iview的组件，如果按照微信的规范，这些组件库会通过miniprogram字段指定自己的构建文件生成目录，开发者工具会把这个目录完全拷贝到最终发布的代码里去，我们就会有两个巨大的组件库占据宝贵的空间。</p>\n<p>我们当然是希望用多少引多少，而不是一股脑全引进去，对，于是MPX提供了按需引用的能力，在下一章<a href=\"#%E6%8C%89%E9%9C%80%E5%BC%95%E7%94%A8\">按需引用</a>细讲。</p>\n<p>以及，组件库目前很少有跨小程序平台的组件库啊，如果我用了vant，支付宝、QQ里没有vant怎么办？也许这是别的框架不怎么推荐使用第三方库的原因，而MPX里，我们帮你把别人的组件库也转了，细节看下下章<a href=\"#%E8%B7%A8%E5%B0%8F%E7%A8%8B%E5%BA%8F%E5%B9%B3%E5%8F%B0\">跨小程序平台</a></p>\n<h3>按需引用</h3>\n<blockquote>\n<p>通过webpack依赖分析收集，使用第三方组件库或者拆分开发大型项目时MPX能保证构建的代码全是要用到的代码</p>\n</blockquote>\n<p>原生小程序本身的编译是遍历项目文件夹里所有的JS，包装成一个AMD包，也就是说项目文件夹里所有的文件，不论是否被使用，都会占用包体积并上传。</p>\n<p>同时，原生微信小程序的npm支持是基于文件夹复制的，第三方包通过声明miniprogram字段指定要拷贝的文件夹，不论使用还是未使用的资源（模板/js/样式/图片），全会被复制到项目文件夹中。</p>\n<p>而我们提供了@mpxjs/webpack-plugin插件，借助webpack生态，解析.mpx文件的json部分或原生的json文件将依赖作为新的入口添加子编译。基于依赖收集，而不是文件遍历。</p>\n<p>带来的好处就是：如果你喜欢vant的按钮，iview的输入框，wux的布局，欢迎尝试MPX，让你能同时使用多个UI框架的同时不用担心应用的体积爆炸。</p>\n<p>同理，面对一个大型项目，我们可以拆成不同的部分，由不同的团队完成后发npm包，在一个主项目中引入即可，具体内容可以看文档<a href=\"https://didi.github.io/mpx/single/json-enhance.html#packages\">JSON增强 - packages</a>一节。</p>\n<p>收集依赖的细节可以查阅文档<a href=\"https://didi.github.io/mpx/understanding/understanding.html#%E7%BC%96%E8%AF%91%E6%9E%84%E5%BB%BA\">编译构建</a>一节。</p>\n<h3>跨小程序平台</h3>\n<blockquote>\n<p>MPX的跨平台方法能带着第三方组件库一起跨小程序平台，同时提供了充足完善的条件编译能力。</p>\n</blockquote>\n<p>在 MPX 1.0 时代，MPX框架是专注提升微信小程序的开发体验，虽然也提供了支付宝版，但代码完全要另写。</p>\n<p>而随着越来越多的 super app 提供了小程序能力，目前至少有5种体系的小程序（微信、支付宝系列、百度系列、头条系列、QQ），如果每一个平台都需要维护一份代码，工程师人数明显不够用了，所以跨小程序平台的能力也是 MPX 2.0 的主打特性。</p>\n<p>我们的跨平台的方法就是转换。都是小程序，语法基本一样，配置、钩子的差异在MPX运行时里提供了抹平。</p>\n<p>而除此之外最大的区别也就是模板上的标签和指令。所以我们实现了一套转换的架子，再编写一份转换规则，即可完成微信小程序到支付宝、百度、头条小程序的转换。</p>\n<p>采用这种转换的模式，非常方便用户理解我们是如何把微信小程序转换成支付宝、百度等小程序平台的。而且只要用户有需求，可以补齐任一套小程序转换其他平台的规则，就可以完成以某个小程序为标准为基础来编写小程序代码以及进一步转换成别的平台的能力。</p>\n<p>再结合前面一直在说的我们对原生小程序的支持，就可以撞出一点不一样的东西，比如，前文提到的第三方组件库跨小程序平台。</p>\n<p>对，我们能帮你把针对微信编写的ui组件库在支付宝、百度上运行起来，带着组件库一起跨小程序平台。</p>\n<p>那么一定会有这样一个问题，就算MPX对原生的支持再怎么牛逼，有的基础能力只有微信平台有，别的平台没有，MPX的转换还能无中生有吗？</p>\n<p>当然不能，其实这个问题对于所有的跨端框架都是一个问题，所以跨端最核心的问题是，如何搞定差异化部分。</p>\n<p>MPX提供了丰富的条件编译能力，可以以文件为维度差别构建，可以以代码块为维度，也可以以代码维度进行差别构建。</p>\n<p>而且MPX的差异化构建能力也是完全基于webpack实现的，所以上面提到的第三方组件库如果确实存在转换不了的地方，比如vant的picker组件使用内联wxs写了一个小方法叫isSimple在模板里调用了，但是这个方法的写法在百度小程序的filter脚本（filter可以理解为百度小程序的wxs）里不支持，因为百度的filter要求必须导出一个对象包裹方法。</p>\n<p>最好的解决办法当然是给vant-weapp提pr帮他们解决一下这个问题，但时间可能会比较慢，所以在MPX里，可以利用webpack的alias能力：</p>\n<p><img src=\"https://dpubstatic.udache.com/static/dpubimg/7f90593b-65b4-40d0-bd93-6396e8a18e64.png\" alt=\"通过alias解决第三方组件的跨平台问题\"></p>\n<p>当尝试构建百度小程序时，会优先去查找pick/index.swan.wxml，再被alias到一个src下的文件，自己修改一下第三方包里有一些小问题的部分即可。</p>\n<p>关于跨平台的条件编译，更多具体信息可以<a href=\"https://didi.github.io/mpx/platform.html#%E8%B7%A8%E5%B9%B3%E5%8F%B0%E7%BC%96%E8%AF%91\">查阅我们的官方文档 - 跨平台条件编译</a></p>\n<h3>能力增强</h3>\n<blockquote>\n<p>通过数据响应、编译时预处理提供了computed/watch，完备的样式类型绑定，双向数据绑定，动态组件等一系列方便开发者更好开发小程序的能力增强</p>\n</blockquote>\n<p>能力增强应该是一个框架提供的最核心最重要的能力了，而MPX也确实在这里下了很大的力气，提供了多且好用的能力增强，不过受限于此处的篇幅，就只简单介绍，细节大家还是<a href=\"https://didi.github.io/mpx/single/template-enhance.html#template%E5%A2%9E%E5%BC%BA%E7%89%B9%E6%80%A7\">查阅我们的文档</a>的好。</p>\n<p>别的框架由于往往基于react/vue的，会给个列表写明不支持哪些能力，用户写的时候习惯使然，往往用了后可能才反应过来哦这个不支持。MPX则是原生的小程序语法写着难受时候突然想起MPX有这个能力。</p>\n<p>列一下MPX增强的能力：</p>\n<ul>\n<li>模板上的增强\n<ul>\n<li>样式类名绑定</li>\n<li>内联事件传参</li>\n<li>动态组件</li>\n<li>双向绑定</li>\n<li>节点获取ref</li>\n</ul>\n</li>\n<li>JS里的增强\n<ul>\n<li>数据响应</li>\n<li>setData优化</li>\n<li>ES6+</li>\n</ul>\n</li>\n<li>样式上的增强\n<ul>\n<li>预处理支持</li>\n<li>rpx转换</li>\n</ul>\n</li>\n<li>JSON里的增强\n<ul>\n<li>packages</li>\n<li>分包资源优化</li>\n</ul>\n</li>\n</ul>\n<p>MPX最显著的能力是数据响应，它衍生出computed/watch，以及双向数据绑定等。这个能力和Vue比较像，不同的是在MPX里是由mobx提供的数据响应能力。</p>\n<p>而同样是数据响应，我们做了一些不一样的优化。</p>\n<h3>性能优势</h3>\n<blockquote>\n<p>通过对模板的解析抽象出访问的数据以保证在提供了数据响应能力的同时不至于劣化性能。</p>\n</blockquote>\n<p>mpvue/wepy/megalo等框架也提供了数据响应的能力，但是数据响应在小程序领域有个较大的问题，微信开发指南里明确提到要注意setData的调用频次和数据量的大小。</p>\n<p>而数据响应最基本的做法就是数据变了就去set数据，这会极大劣化小程序的性能表现。</p>\n<p>而MPX通过对模板进行解析，抽象出对应的render函数，在调用setData发送数据前执行render函数找到真正需要发送的数据。</p>\n<p>效果如图：</p>\n<p><img src=\"https://dpubstatic.udache.com/static/dpubimg/9399b75c-5c81-4584-89b6-5e7f2c2ba7ba.png\" alt=\"小程序性能分析\"></p>\n<p>小程序开发者工具的audits面板能辅助用户分析出可能需要优化的点。正如前文所说，MPX在红框部分，尤其是红框里的第三条，不将模板上未使用的数据发送到渲染层上做了极大的优化。</p>\n<p>只要不出现渲染函数执行失败（会有warning在console里提示，同时兜底逻辑会进行全量setData以保证程序仍可正常运行），使用MPX开发的小程序就永远不用担心发送了模板未使用的数据。</p>\n<blockquote>\n<p>为了不降低首次渲染的速度，我们未对构造器里声明的data初始值做这个分析，所以也不要因为MPX有这个特性就大肆在data上声明过多不在模板上使用的数据。</p>\n</blockquote>\n<p>虽然只是一个小小的TODO MVC示例，但是这个优化和应用的规模没关系，而且同时大家可以尝试别家的小demo对比看看。</p>\n<p>这个优化的细节可以看<a href=\"https://didi.github.io/mpx/articles/2.0.html\">前一篇文章</a>，或者我们的文档<a href=\"https://didi.github.io/mpx/understanding/understanding.html#%E6%95%B0%E6%8D%AE%E5%93%8D%E5%BA%94%E4%B8%8E%E6%80%A7%E8%83%BD%E4%BC%98%E5%8C%96\">MPX运行机制 - 数据响应与性能优化</a></p>\n<h2>总结</h2>\n<p>与目前市面上的诸多框架相比，MPX希望以原生小程序为基础，全面拥抱原生小程序，在原生小程序的基础上做增强，通过尽可能少的转换实现尽可能多的能力增强，在提升小程序开发体验的同时，保证不因转换或框架的问题产生过多的坑。</p>\n<p>MPX框架的目标用户是对小程序质量有较高要求的开发者，如果你是原生小程序开发者，或者厌倦了解决某些以web框架DSL语法为基础的转换框架造成的坑，欢迎尝试MPX框架。</p>\n<p><a href=\"https://github.com/didi/mpx\">MPX GITHUB：https://github.com/didi/mpx</a></p>\n","date_published":"2019-05-26T00:00:00.000Z","tags":["MPX"],"language":"zh"},{"id":"https://www.lihuanyu.com/posts/2017/docker-for-windows%E4%B8%8D%E5%93%8D%E5%BA%94react%E9%A1%B9%E7%9B%AE%E6%94%B9%E5%8F%98%E5%90%8E%E7%9A%84%E9%87%8D%E7%BC%96%E8%AF%91/","url":"https://www.lihuanyu.com/posts/2017/docker-for-windows%E4%B8%8D%E5%93%8D%E5%BA%94react%E9%A1%B9%E7%9B%AE%E6%94%B9%E5%8F%98%E5%90%8E%E7%9A%84%E9%87%8D%E7%BC%96%E8%AF%91/","title":"Docker for Windows 下的前端热更新问题","summary":"已并入《重新认识 Docker：开发环境、Linux 性能开销与 Redis 实战》。","content_html":"<p>本文记录的是 Docker for Windows 早期文件监听不稳定导致 React 项目无法触发重新编译的问题。当时通过额外 watcher 规避了这个问题，但它更适合作为历史经验，而不是今天的默认方案。</p>\n<p>完整讨论见：</p>\n<p><a href=\"/posts/2025/%E9%87%8D%E6%96%B0%E8%AE%A4%E8%AF%86Docker%E7%9A%84%E6%80%A7%E8%83%BD%E5%BC%80%E9%94%80/\">重新认识 Docker：开发环境、Linux 性能开销与 Redis 实战</a></p>\n<p>今天在 Windows 上做容器化开发，更推荐使用 WSL2，并把项目放在 Linux 发行版的文件系统里。对前端热更新、依赖安装和大量小文件读写来说，宿主机文件系统与容器之间的边界仍然是需要重点关注的性能点。</p>\n","date_published":"2017-12-05T00:00:00.000Z","date_modified":"2026-05-04T00:00:00.000Z","tags":["docker","windows","前端"],"language":"zh"},{"id":"https://www.lihuanyu.com/posts/2017/%E5%89%8D%E7%AB%AF%E9%A1%B9%E7%9B%AE%E5%B7%A5%E7%A8%8B%E5%8C%96%E5%AE%9E%E8%B7%B5/","url":"https://www.lihuanyu.com/posts/2017/%E5%89%8D%E7%AB%AF%E9%A1%B9%E7%9B%AE%E5%B7%A5%E7%A8%8B%E5%8C%96%E5%AE%9E%E8%B7%B5/","title":"前端项目工程化实践","summary":"以 Antue 组件库为例，记录前端开源项目接入 Travis CI、lint、测试、构建和 GitHub Pages 自动部署的工程化实践。","content_html":"<p>17年9月起和朋友合作了一个项目，<a href=\"https://github.com/zzuu666/antue\">一套组件库Antue</a>，好听点说叫造轮子。主要是把蚂蚁金服的Ant Design给“翻译&quot;成Vue可用的组件库。这是一个蛮正式的项目，规模也挺大，所以给了我一个实践工程化的好场景。</p>\n<h2>什么是前端工程化</h2>\n<p><a href=\"https://github.com/ruanyf/jstraining/blob/master/docs/engineering.md\">阮一峰 - 前端工程简介</a></p>\n<p>我认为应该是包含包管理、代码构建、代码lint、单元测试、持续集成等等的一个合集。</p>\n<p>包管理、代码构建、lint、单测在目前的前端项目中已经是标配了，用vue-cli生成的模板天然地为大家处理好了这些问题。所以我的实践主要是最后一步，持续集成。</p>\n<p>持续集成的流程、概念、好处等等在上面的阮一峰的文章中说得很明确了。</p>\n<h2>为什么需要</h2>\n<p>很多朋友在业务项目中很少使用持续集成，我也是，主要是业务上使用持续集成需要考虑成本是否划算，可能是做的业务都不太好吧，如果是一个公司的主航道业务，需要一直一直持续迭代持续开发，持续集成会是一个非常好的助手我猜。</p>\n<p>而开源的，让别人使用的lib级别的东西，再怎么严谨都是不为过的，更何况，一个优秀的持续集成工作流，能保证项目的质量，让迭代成本、维护成本变低，是一件双赢的好事。</p>\n<p>有了持续集成机制，我们可以保证每次对主干的合并后，能检查以下项：</p>\n<ul>\n<li>开发者提交的代码是否遵循了eslint的规则，保证风格一致，无低级错误</li>\n<li>开发者提交的代码是否能够通过单元测试，避免改动或重构影响其他相关的地方</li>\n<li>开发者提交的代码是否能够正常build，给出一个正确的压缩包</li>\n</ul>\n<h2>如何做</h2>\n<p>因为这是一个github开源项目，Travis是一个非常优秀、非常方便的选择。</p>\n<p>仅需要项目owner用github账户登录一下Travis，勾选需要启用的对应的项目。然后编写Travis的配置文件：.travis.yml即可。</p>\n<p>配置内容也比较简单，如下：</p>\n<pre><code class=\"language-yaml\">language: node_js\nnode_js:\n  - 8.5.0\nbefore_install:\n  - npm install\nscript:\n  - npm run lint\n  - npm test\n  - npm run build\nbefore_deploy: \n  - node scripts/generate.js -a\n  - node scripts/generate.js -r\n  - npm run build:site\ndeploy:\n  provider: pages\n  skip_cleanup: true\n  github_token: $GH_TOKEN\n  local_dir: site/dist\n  on:\n    branch: master\n</code></pre>\n<p>可以理解成生命周期的各个钩子吧，先<code>npm install</code>安装依赖，再运行正式的script：先lint，再跑单测，最后试试构建。</p>\n<p>任一环节出错就会发邮件告知项目相关人员，可以通过查看CI的log信息来检查到底是什么问题，在本地复现、修复。</p>\n<p>这里有一个关键问题是，如何保证大家认可Travis的结果，有很多开源项目都使用了Travis，但是我也见过不少Travis报着错，还在继续merge到master分支的项目。</p>\n<p>其实也很好解决，只是个理念的问题，相信制度、流程，而不是人的自觉，github对项目的设定中提供了保护分支的选项，通过把master设为保护分支，可以要求pull request必须通过CI才可以合入，我们的项目还更进一步，加了一条必须有合作者的code review。</p>\n<h2>效果如何？</h2>\n<p>首先，antue的master分支的component文件夹下的组件代码，绝对没有不符合standard规范的JS代码。</p>\n<p>其次，未来的重构中，不会有因重构导致有完善单元测试的组件出bug（很惭愧，我写的组件还没写单元测试）。</p>\n<p>最后，任何一台安装了合适版本node的能联网的电脑，都可以完美编译该项目（也许不能，毕竟目前的开发者使用的都是mac，可能会有平台兼容性问题，吹个牛逼也不犯法不是？）。很多人的Travis配置中可能只有跑了一个单元测试，但构建这一点其实蛮重要的，给人一个能编译（构建）过的项目，有问题比不能编译过的项目会好查太多太多。</p>\n<p>同时，有了Travis的构建通过邮件，大家能很开心很自信地往下写。</p>\n<p>所以有兴趣的同学，欢迎参与开发这个组件库，我们一点也不担心不同开发人员的不同风格是否会导致项目变得奇怪。</p>\n<h2>还可以做什么？</h2>\n<p>看上面的配置文件，最下面的<code>deploy</code>说明我们用Travis进行的自动部署。</p>\n<p>这个需求主要是这样的，主owner希望对antd进行像素级复制，所以文档网站也是用类似于andt的手法，通过markdown生成的（没有antd的毕昇系统那么叼，但也显得很专业嘿嘿）。</p>\n<p>发布的方式是手工执行命令：</p>\n<pre><code class=\"language-bash\">node scripts/generate.js -a\nnode scripts/generate.js -r\nnpm run build:site\ngit checkout gh-pages\ncp ./site/dist/* ./\ngit add .\ngit commit -m '第XX次 update doc'\ngit push\n</code></pre>\n<p>可以看出这个部署到github pages的操作是非常的固定的，且我们又知道我们的master分支是随时处于一个待发布的OK的状态的，那么我们为什么不让每次merge到master后就自动部署到pages上去呢？</p>\n<p>一开始我想的是通过自己编写一个bash脚本，让Travis每次来执行这个脚本，其实这个方案也是可以的，只是不好调试，费了半天劲没有解决只master版本才部署这个问题，就很伤。</p>\n<p>查了半天Travis的文档后发现它内置的部署provider里本身就有pages这种方案（Travis文档做得挺好，就是搜索功能非常不好用），就改成目前的这样。</p>\n<p>自动部署的效果很好，每完成一个新组件，合并master后，就可以立即看到效果，想参与github开源项目的同学走过路过不要错过，提交就可以拿着<a href=\"https://zzuu666.github.io/antue/\">这个网页</a>找到自己写的组件去和人吹牛逼啦！</p>\n<p>本项目在持续集成上的实践就酱紫啦，要是还有什么需要做的好点子欢迎和我分享。</p>\n","date_published":"2017-11-30T00:00:00.000Z","tags":[],"language":"zh"},{"id":"https://www.lihuanyu.com/en/posts/2017/cors-cookies-credentials-samesite/","url":"https://www.lihuanyu.com/en/posts/2017/cors-cookies-credentials-samesite/","title":"CORS Cookies Not Sent? Fix Credentials and SameSite","summary":"Cross-origin cookies not sent? Check credentials, explicit CORS response headers, Set-Cookie ownership, SameSite, Secure, and browser cookie policy.","content_html":"<p>A cross-origin cookie reaches an API only when the request, CORS response, cookie attributes, and browser policy all allow it. Start with <code>credentials: 'include'</code>, an explicit <code>Access-Control-Allow-Origin</code>, <code>Access-Control-Allow-Credentials: true</code>, and a cookie set by the API through <code>Set-Cookie</code>.</p>\n<p>If Chrome DevTools shows a cookie but the next request omits it, use the checks below before changing application code.</p>\n<p><a href=\"/posts/2017/%E8%B0%88%E8%B0%88CORS%E4%B8%8B%E5%89%8D%E7%AB%AF%E7%9A%84cookie/\">Chinese version of this article</a></p>\n<h2>Fix CORS cookies in five checks</h2>\n<p>The browser evaluates several independent rules before it sends a cookie:</p>\n<div class=\"table-scroll\"><table>\n<thead>\n<tr>\n<th>Check</th>\n<th>Required configuration</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Request credentials</td>\n<td>Use <code>credentials: 'include'</code> with <code>fetch</code>, or <code>withCredentials: true</code> with XHR and Axios</td>\n</tr>\n<tr>\n<td>Allowed origin</td>\n<td>Return the requesting origin in <code>Access-Control-Allow-Origin</code>; credentialed responses cannot use <code>*</code></td>\n</tr>\n<tr>\n<td>Credential permission</td>\n<td>Return <code>Access-Control-Allow-Credentials: true</code></td>\n</tr>\n<tr>\n<td>Cookie ownership</td>\n<td>Let the API set its cookie through <code>Set-Cookie</code>; frontend JavaScript cannot set a cookie for an unrelated API host</td>\n</tr>\n<tr>\n<td>Cookie policy</td>\n<td>Match <code>Domain</code>, <code>Path</code>, <code>SameSite</code>, and <code>Secure</code>, then check whether the browser blocks third-party cookies</td>\n</tr>\n</tbody>\n</table>\n</div><p>These checks solve different parts of the request. <code>withCredentials</code> cannot repair the wrong cookie domain, and CORS headers cannot override browser privacy policy.</p>\n<h2>Separate same-origin, same-site, and cookie scope</h2>\n<p>Three concepts are often mixed together:</p>\n<ul>\n<li><strong>Same-origin</strong>: scheme, host, and port are all the same. <code>https://www.example.com</code> and <code>https://api.example.com</code> are different origins. So are <code>http://localhost:3000</code> and <code>http://localhost:63342</code>.</li>\n<li><strong>Same-site</strong>: usually based on scheme plus the registrable domain. <code>https://www.example.com</code> and <code>https://api.example.com</code> are usually same-site. <code>https://example.com</code> and <code>https://other.com</code> are cross-site.</li>\n<li><strong>Cookie scope</strong>: decided by the host that set the cookie plus attributes such as <code>Domain</code> and <code>Path</code>. Ports are not part of cookie scope.</li>\n</ul>\n<p>CORS controls whether a script from one origin can read a response from another origin. Cookies control which stored cookies are automatically attached to a request for a host/path. SameSite controls whether cookies are allowed on same-site or cross-site requests.</p>\n<p>Because these are different checks, some cases look surprising:</p>\n<ul>\n<li><code>localhost:63342</code> requesting <code>localhost:3000</code>: cross-origin, so CORS is needed; but both use the <code>localhost</code> host, so cookies can look shared while debugging.</li>\n<li><code>www.example.com</code> requesting <code>api.example.com</code>: cross-origin, so CORS is needed; but usually same-site, so <code>SameSite=Lax</code> may still allow cookies.</li>\n<li><code>app.example.com</code> requesting <code>api.other.com</code>: cross-origin and cross-site, so CORS, SameSite, and third-party cookie policy all matter.</li>\n</ul>\n<h2>Configure a credentialed CORS request</h2>\n<p>The frontend must explicitly include credentials. Fetch only sends cookies by default for same-origin requests. Cross-origin requests need <code>credentials: 'include'</code>:</p>\n<pre><code class=\"language-js\">await fetch('https://api.example.com/me', {\n  method: 'GET',\n  credentials: 'include'\n});\n</code></pre>\n<p>For XHR or axios, the corresponding setting is:</p>\n<pre><code class=\"language-js\">xhr.withCredentials = true;\n</code></pre>\n<pre><code class=\"language-js\">axios.get('https://api.example.com/me', {\n  withCredentials: true\n});\n</code></pre>\n<p>The server must also allow credentialed CORS requests:</p>\n<ul>\n<li><code>Access-Control-Allow-Origin</code> must be an explicit origin, not <code>*</code>.</li>\n<li><code>Access-Control-Allow-Credentials</code> must be <code>true</code>.</li>\n<li>If the server reflects allowed origins dynamically, it should also return <code>Vary: Origin</code> to avoid cache confusion.</li>\n<li>Preflight <code>OPTIONS</code> requests do not include cookies, but their responses still need to indicate whether the real request is allowed to include credentials.</li>\n</ul>\n<p>An Express example:</p>\n<pre><code class=\"language-js\">const allowList = new Set([\n  'https://www.example.com'\n]);\n\napp.use((req, res, next) =&gt; {\n  const origin = req.headers.origin;\n\n  if (allowList.has(origin)) {\n    res.setHeader('Access-Control-Allow-Origin', origin);\n    res.setHeader('Vary', 'Origin');\n    res.setHeader('Access-Control-Allow-Credentials', 'true');\n    res.setHeader('Access-Control-Allow-Methods', 'GET,POST,PUT,DELETE,OPTIONS');\n    res.setHeader('Access-Control-Allow-Headers', 'Content-Type, Authorization');\n  }\n\n  if (req.method === 'OPTIONS') {\n    return res.sendStatus(204);\n  }\n\n  next();\n});\n</code></pre>\n<p>Finally, the cookie should be set by the API host through <code>Set-Cookie</code>. If the page is <code>https://www.example.com</code> and the API is <code>https://api.example.com</code>, they are same-site but cross-origin:</p>\n<pre><code class=\"language-http\">Set-Cookie: __Host-sid=...; Path=/; HttpOnly; Secure; SameSite=Lax\n</code></pre>\n<p>This is a host-only cookie. It is sent only to <code>api.example.com</code>. <code>HttpOnly</code> prevents JavaScript from reading it, which is appropriate for session cookies. <code>Secure</code> requires HTTPS. <code>SameSite=Lax</code> is often enough for same-site requests.</p>\n<p>If the page and API are cross-site, for example <code>https://app.example.com</code> calling <code>https://api.other.com</code>, a cookie intended for cross-site requests needs at least:</p>\n<pre><code class=\"language-http\">Set-Cookie: sid=...; Path=/; HttpOnly; Secure; SameSite=None\n</code></pre>\n<p>That only means the cookie has attributes that allow cross-site sending. It does not guarantee the cookie will work, because browser or user-level third-party cookie policy may still block it.</p>\n<h2>Why document.cookie does not fix it</h2>\n<p><code>document.cookie = 'sid=123'</code> writes a cookie for the current page’s host. If the page is on <code>www.example.com</code>, the script cannot write a cookie for <code>api.other.com</code>.</p>\n<p>Even with the <code>Domain</code> attribute, a page can only set cookies for the current host or a parent domain that contains it. For example, a response from <code>api.example.com</code> can set <code>Domain=example.com</code>, but it cannot set <code>Domain=other.com</code>.</p>\n<p>That is why returning a cookie value in JSON and asking the frontend to write it into <code>document.cookie</code> is usually the wrong design:</p>\n<ul>\n<li>The cookie belongs to the frontend page’s domain, not the API domain.</li>\n<li>If the session cookie needs <code>HttpOnly</code>, JavaScript should not be able to write or read it.</li>\n<li><code>Set-Cookie</code> is a forbidden response header for frontend JavaScript. <code>Access-Control-Expose-Headers: Set-Cookie</code> does not make it readable.</li>\n</ul>\n<p>The correct flow is: the API returns <code>Set-Cookie</code>; the browser stores it if CORS, credentials, cookie attributes, and browser policy all allow it; later requests attach the cookie automatically.</p>\n<h2>Account for modern SameSite defaults</h2>\n<p>Older CORS articles often said that cross-origin cookies work once <code>withCredentials</code> and <code>Access-Control-Allow-Credentials</code> are configured. Today, that advice is missing SameSite.</p>\n<p>Modern browsers usually treat cookies without an explicit SameSite value as <code>Lax</code>. <code>Lax</code> sends cookies on same-site requests and on some top-level cross-site navigations, but it does not freely attach cookies to cross-site <code>fetch</code>, XHR, or iframe subresource requests.</p>\n<p>So if a cross-site API request depends on cookies, the cookie generally needs:</p>\n<pre><code class=\"language-http\">Set-Cookie: sid=...; Path=/; HttpOnly; Secure; SameSite=None\n</code></pre>\n<p>Two details matter:</p>\n<ul>\n<li><code>SameSite=None</code> must be paired with <code>Secure</code>.</li>\n<li><code>Secure</code> means production should use HTTPS. <code>localhost</code> has development exceptions, but local behavior should not be treated as production behavior.</li>\n</ul>\n<p>If the frontend and API can be placed under the same site, prefer that design:</p>\n<ul>\n<li><code>https://www.example.com</code></li>\n<li><code>https://api.example.com</code></li>\n</ul>\n<p>This still requires CORS because the origins are different, but the SameSite pressure is much lower because the request is usually same-site.</p>\n<h2>CORS cannot bypass third-party cookie policy</h2>\n<p>Even if CORS headers, <code>credentials: 'include'</code>, and <code>SameSite=None; Secure</code> are all correct, cookies can still be blocked. Browser third-party cookie policy sits outside the CORS configuration.</p>\n<p>MDN’s CORS documentation explicitly notes that credentialed cross-origin requests are still subject to third-party cookie policies. Frontend and server settings cannot override user-agent policy.</p>\n<p>At minimum, browser behavior should be understood separately:</p>\n<ul>\n<li>Safari/WebKit has restricted third-party cookies for a long time and moved to full third-party cookie blocking in 2020.</li>\n<li>Firefox Enhanced Tracking Protection blocks some tracking-related third-party cookies.</li>\n<li>Chrome announced in 2025 that it would keep giving users control over third-party cookies instead of launching a new standalone prompt. Incognito mode blocks third-party cookies by default, and users can also disable them in privacy settings.</li>\n</ul>\n<p>So third-party cookies should not be treated as stable login infrastructure for normal web applications. The more robust design is to avoid putting the frontend site and login cookie site on completely different sites.</p>\n<p>If the product is truly a third-party embedded component, such as an iframe widget, map, support chat, payment flow, or cross-site embedded app, then APIs such as Storage Access API and CHIPS/Partitioned Cookies may be worth evaluating. They have specific use cases and should not be the default for ordinary frontend/backend login.</p>\n<h2>Choose a stable deployment design</h2>\n<p>In order of stability, I would choose:</p>\n<ol>\n<li><strong>Same-origin deployment</strong>: serve the frontend and API under the same origin, or use Nginx/BFF to proxy <code>/api</code> to the backend. Cookies are simplest and CORS mostly disappears.</li>\n<li><strong>Same-site but cross-origin</strong>: for example <code>www.example.com</code> plus <code>api.example.com</code>. CORS and <code>credentials: 'include'</code> are still needed, but cookies remain in the same-site model.</li>\n<li><strong>Cross-site without cookie-based API identity</strong>: public APIs, cross-organization APIs, and mobile APIs are better served by OAuth, short-lived tokens, or <code>Authorization</code> headers.</li>\n<li><strong>Cross-site and cookie-based</strong>: only choose this when browser compatibility, user settings, and embedded context are fully understood, and when there is a fallback for blocked third-party cookies.</li>\n</ol>\n<p>A reverse proxy is not a primitive workaround. For applications you control, making the browser see the frontend and API as one site is often more reliable than fighting browser privacy policy.</p>\n<h2>Protect the cookie security boundary</h2>\n<p>Cookies are attached automatically by the browser. That is also why CSRF exists. CORS is not CSRF protection. A cross-site form submission or simple request can still be sent; the attacker page may simply be unable to read the response.</p>\n<p>If an API uses cookies as login state, at least consider:</p>\n<ul>\n<li>Use <code>HttpOnly; Secure</code> for session cookies.</li>\n<li>Prefer <code>SameSite=Lax</code> when possible instead of <code>SameSite=None</code>.</li>\n<li>For state-changing requests, validate a CSRF token or check request-origin signals such as <code>Origin</code> and <code>Sec-Fetch-Site</code>.</li>\n<li>Do not reflect every <code>Origin</code> into <code>Access-Control-Allow-Origin</code>.</li>\n<li>Do not use <code>Access-Control-Allow-Origin: *</code> on credentialed CORS responses.</li>\n</ul>\n<p>Cookies carry identity automatically. They do not prove that a request is trustworthy.</p>\n<h2>Debug CORS cookies in order</h2>\n<p>When a CORS cookie is not being received, check in this order:</p>\n<ol>\n<li>Is the request cross-origin? Compare scheme, host, and port.</li>\n<li>Did the frontend set <code>fetch(..., { credentials: 'include' })</code> or <code>withCredentials = true</code>?</li>\n<li>Does the response contain an explicit <code>Access-Control-Allow-Origin</code>, and is it not <code>*</code>?</li>\n<li>Does the response contain <code>Access-Control-Allow-Credentials: true</code>?</li>\n<li>If origin is dynamic, is <code>Vary: Origin</code> present?</li>\n<li>Is the cookie set by the API host through <code>Set-Cookie</code>, not returned in the response body?</li>\n<li>Do <code>Domain</code> and <code>Path</code> cover the next request URL?</li>\n<li>For cross-site requests, is the cookie <code>SameSite=None; Secure</code>?</li>\n<li>Is production using HTTPS?</li>\n<li>Is the browser blocking third-party cookies?</li>\n<li>Does DevTools show the cookie as blocked, and what is the blocked reason?</li>\n<li>Does the Application panel show the cookie under the expected site?</li>\n</ol>\n<p>In Chrome DevTools, the <code>Cookies</code> subpanel inside a specific Network request is often more useful than looking only at raw headers. Cookies blocked by SameSite, Secure, Domain, or third-party cookie policy often show a reason there or in the Issues panel.</p>\n<h2>What must line up</h2>\n<p>CORS cookies work only when the whole chain lines up:</p>\n<ul>\n<li>The frontend allows credentials.</li>\n<li>The server allows that specific origin to send credentials.</li>\n<li>The cookie is set by the target domain through <code>Set-Cookie</code>.</li>\n<li>Domain, Path, SameSite, and Secure match the request scenario.</li>\n<li>Browser third-party cookie policy does not block it.</li>\n</ul>\n<p>The root cause in that old 2017 bug was: the frontend cannot write cookies for the backend domain. The modern addition is: even correctly set backend cookies must be designed with SameSite, Secure, and third-party cookie restrictions in mind.</p>\n<h2>Further reading</h2>\n<ul>\n<li><a href=\"https://developer.mozilla.org/en-US/docs/Web/HTTP/Guides/CORS\">MDN: Cross-Origin Resource Sharing</a></li>\n<li><a href=\"https://developer.mozilla.org/en-US/docs/Web/HTTP/Reference/Headers/Set-Cookie\">MDN: Set-Cookie</a></li>\n<li><a href=\"https://developer.mozilla.org/en-US/docs/Web/HTTP/Guides/Cookies\">MDN: Using HTTP cookies</a></li>\n<li><a href=\"https://developer.mozilla.org/en-US/docs/Web/API/Fetch_API/Using_Fetch#including_credentials\">MDN: Using the Fetch API - Including credentials</a></li>\n<li><a href=\"https://developer.chrome.com/docs/devtools/application/cookies/\">Chrome for Developers: View, add, edit, and delete cookies</a></li>\n<li><a href=\"https://privacysandbox.com/news/privacy-sandbox-next-steps/\">Privacy Sandbox: Next steps for Privacy Sandbox and tracking protections in Chrome</a></li>\n<li><a href=\"https://webkit.org/blog/10218/full-third-party-cookie-blocking-and-more/\">WebKit: Full Third-Party Cookie Blocking and More</a></li>\n<li><a href=\"https://developer.mozilla.org/en-US/docs/Web/Privacy/Privacy_sandbox/Partitioned_cookies\">MDN: Cookies Having Independent Partitioned State</a></li>\n</ul>\n","date_published":"2017-09-02T00:00:00.000Z","date_modified":"2026-08-01T00:00:00.000Z","tags":["CORS","Cookie","SameSite","Frontend","Browser"],"language":"en"},{"id":"https://www.lihuanyu.com/posts/2017/%E8%B0%88%E8%B0%88CORS%E4%B8%8B%E5%89%8D%E7%AB%AF%E7%9A%84cookie/","url":"https://www.lihuanyu.com/posts/2017/%E8%B0%88%E8%B0%88CORS%E4%B8%8B%E5%89%8D%E7%AB%AF%E7%9A%84cookie/","title":"CORS 下 Cookie 为什么收不到：从 withCredentials 到 SameSite","summary":"CORS 下 Cookie 能不能生效，不只取决于 withCredentials，还取决于服务端 CORS 头、Cookie Domain、SameSite、Secure 和浏览器第三方 Cookie 策略。","content_html":"<p>很多年前排查过一个问题：前端页面通过 CORS 请求后端，后端希望前端把响应里的某个字段写入 <code>document.cookie</code>，再在下一次请求里带回去。前端确实写了 Cookie，Chrome DevTools 里也能看到，但后端就是收不到。</p>\n<p>当时的结论很简单：Cookie 是按域存储和发送的，页面脚本写下的 Cookie 属于当前页面所在的域，不会因为一次跨域请求就变成接口域名的 Cookie。后端应该用 <code>Set-Cookie</code>，而不是把 Cookie 值塞进 response body 让前端手工写。</p>\n<p><a href=\"/en/posts/2017/cors-cookies-credentials-samesite/\">English version: Why Cookies Fail in CORS: From withCredentials to SameSite</a></p>\n<p>这个结论今天仍然成立，但已经不够完整。现代浏览器补上了 SameSite 默认值、<code>SameSite=None</code> 必须搭配 <code>Secure</code>、第三方 Cookie 限制、分区 Cookie 等机制。现在排查 CORS 下 Cookie 收不到，不能只盯着 <code>withCredentials</code>。</p>\n<h2>先区分三个概念</h2>\n<p>讨论 CORS 和 Cookie 时，最容易混在一起的是这三个概念：</p>\n<ul>\n<li><strong>同源</strong>：scheme、host、port 都相同。<code>https://www.example.com</code> 和 <code>https://api.example.com</code> 不同源；<code>http://localhost:3000</code> 和 <code>http://localhost:63342</code> 也不同源。</li>\n<li><strong>同站</strong>：通常看 scheme 加可注册域名。<code>https://www.example.com</code> 和 <code>https://api.example.com</code> 通常同站；<code>https://example.com</code> 和 <code>https://other.com</code> 不同站。</li>\n<li><strong>Cookie 作用域</strong>：由设置 Cookie 的 host、<code>Domain</code>、<code>Path</code> 等属性决定。端口不是 Cookie 作用域的一部分。</li>\n</ul>\n<p>CORS 管的是“一个 origin 的脚本能不能读取另一个 origin 的响应”。Cookie 管的是“请求某个 host/path 时，哪些 Cookie 会自动带上”。SameSite 管的是“当前请求是不是跨站，跨站时 Cookie 还能不能带”。</p>\n<p>这三个判断维度不同，所以会出现一些看起来反直觉的场景：</p>\n<ul>\n<li><code>localhost:63342</code> 请求 <code>localhost:3000</code>：不同源，需要 CORS；但 Cookie 的 host 都是 <code>localhost</code>，调试时容易在同一个 Cookie 面板里看到。</li>\n<li><code>www.example.com</code> 请求 <code>api.example.com</code>：不同源，需要 CORS；但通常同站，<code>SameSite=Lax</code> 不一定会拦住 Cookie。</li>\n<li><code>app.example.com</code> 请求 <code>api.other.com</code>：不同源且不同站，既要 CORS，也会受到 SameSite 和第三方 Cookie 策略影响。</li>\n</ul>\n<h2>一条能工作的 CORS Cookie 链路</h2>\n<p>前端要明确带凭据。Fetch 默认只在同源请求里带 Cookie，跨源请求需要 <code>credentials: 'include'</code>：</p>\n<pre><code class=\"language-js\">await fetch('https://api.example.com/me', {\n  method: 'GET',\n  credentials: 'include'\n});\n</code></pre>\n<p>如果使用 XHR 或 axios，对应的是：</p>\n<pre><code class=\"language-js\">xhr.withCredentials = true;\n</code></pre>\n<pre><code class=\"language-js\">axios.get('https://api.example.com/me', {\n  withCredentials: true\n});\n</code></pre>\n<p>服务端也要明确允许带凭据。关键点是：</p>\n<ul>\n<li><code>Access-Control-Allow-Origin</code> 必须是明确的 origin，不能是 <code>*</code>。</li>\n<li><code>Access-Control-Allow-Credentials</code> 必须是 <code>true</code>。</li>\n<li>如果按请求的 <code>Origin</code> 动态返回 <code>Access-Control-Allow-Origin</code>，要加 <code>Vary: Origin</code>，避免缓存污染。</li>\n<li>预检请求 <code>OPTIONS</code> 不会带 Cookie，但预检响应仍要告诉浏览器后续真实请求是否允许带凭据。</li>\n</ul>\n<p>一个 Express 示例：</p>\n<pre><code class=\"language-js\">const allowList = new Set([\n  'https://www.example.com'\n]);\n\napp.use((req, res, next) =&gt; {\n  const origin = req.headers.origin;\n\n  if (allowList.has(origin)) {\n    res.setHeader('Access-Control-Allow-Origin', origin);\n    res.setHeader('Vary', 'Origin');\n    res.setHeader('Access-Control-Allow-Credentials', 'true');\n    res.setHeader('Access-Control-Allow-Methods', 'GET,POST,PUT,DELETE,OPTIONS');\n    res.setHeader('Access-Control-Allow-Headers', 'Content-Type, Authorization');\n  }\n\n  if (req.method === 'OPTIONS') {\n    return res.sendStatus(204);\n  }\n\n  next();\n});\n</code></pre>\n<p>最后，Cookie 应该由接口域名通过 <code>Set-Cookie</code> 设置。比如页面在 <code>https://www.example.com</code>，接口在 <code>https://api.example.com</code>，两者同站但不同源：</p>\n<pre><code class=\"language-http\">Set-Cookie: __Host-sid=...; Path=/; HttpOnly; Secure; SameSite=Lax\n</code></pre>\n<p>这里的 Cookie 是 host-only Cookie，只会发给 <code>api.example.com</code>。<code>HttpOnly</code> 让前端脚本无法读取它，适合会话 Cookie；<code>Secure</code> 要求 HTTPS；<code>SameSite=Lax</code> 在同站请求中通常足够。</p>\n<p>如果页面和接口是不同站，比如 <code>https://app.example.com</code> 请求 <code>https://api.other.com</code>，想让 Cookie 参与跨站请求，Cookie 至少要这样：</p>\n<pre><code class=\"language-http\">Set-Cookie: sid=...; Path=/; HttpOnly; Secure; SameSite=None\n</code></pre>\n<p>但这只表示 Cookie 具备跨站发送的属性，不代表一定能用。浏览器或用户设置仍可能阻止第三方 Cookie。</p>\n<h2>为什么前端写 Cookie 后后端收不到</h2>\n<p><code>document.cookie = 'sid=123'</code> 写的是当前页面所在 host 的 Cookie。页面在 <code>www.example.com</code>，脚本就不能给 <code>api.other.com</code> 写 Cookie。</p>\n<p>即使用 <code>Domain</code>，也只能设置当前 host 或它的父域，不能设置任意外部域。比如从 <code>api.example.com</code> 可以设置 <code>Domain=example.com</code>，让 Cookie 覆盖同一可注册域名下的子域；但不能设置 <code>Domain=other.com</code>。</p>\n<p>这也是为什么“后端把 Cookie 值放在 JSON 里，让前端写到 <code>document.cookie</code>”通常是错误方案：</p>\n<ul>\n<li>写出来的是前端页面域名的 Cookie，不是接口域名的 Cookie。</li>\n<li>如果会话 Cookie 需要 <code>HttpOnly</code>，前端脚本本来就不应该能写。</li>\n<li><code>Set-Cookie</code> 是浏览器特殊处理的响应头，前端 JavaScript 不能读取它；即使服务端加 <code>Access-Control-Expose-Headers: Set-Cookie</code> 也没用。</li>\n</ul>\n<p>正确链路应该是：接口响应里返回 <code>Set-Cookie</code>，浏览器在符合 CORS、credentials、Cookie 属性和浏览器策略的前提下自动保存；后续请求再由浏览器自动带上。</p>\n<h2>SameSite 改变了很多旧经验</h2>\n<p>早期很多文章会说：CORS 配好 <code>withCredentials</code> 和 <code>Access-Control-Allow-Credentials</code>，跨域 Cookie 就能正常用。今天这句话少了 SameSite。</p>\n<p>现代浏览器通常把未声明 SameSite 的 Cookie 当成 <code>Lax</code>。<code>Lax</code> 会在同站请求中发送，也会在用户进行顶层导航的部分跨站场景中发送，但不会为了普通跨站 <code>fetch</code>、XHR、iframe 子资源请求随便发送。</p>\n<p>因此，跨站接口请求如果依赖 Cookie，一般需要：</p>\n<pre><code class=\"language-http\">Set-Cookie: sid=...; Path=/; HttpOnly; Secure; SameSite=None\n</code></pre>\n<p>注意两个细节：</p>\n<ul>\n<li><code>SameSite=None</code> 必须搭配 <code>Secure</code>。</li>\n<li><code>Secure</code> 意味着生产环境必须使用 HTTPS；<code>localhost</code> 是调试例外，但不要把本地表现直接当成线上表现。</li>\n</ul>\n<p>如果前端和 API 只是不同子域，优先把它们放在同一个站点下，例如：</p>\n<ul>\n<li><code>https://www.example.com</code></li>\n<li><code>https://api.example.com</code></li>\n</ul>\n<p>这种架构仍然需要 CORS，因为它们不同源；但 SameSite 压力会小很多，因为它们通常是同站请求。</p>\n<h2>第三方 Cookie 策略不能靠 CORS 绕过</h2>\n<p>CORS 头、<code>credentials: 'include'</code>、<code>SameSite=None; Secure</code> 都配对了，也仍然可能收不到 Cookie。原因是浏览器的第三方 Cookie 策略还在更外层。</p>\n<p>MDN 在 CORS 文档里也明确提醒：带凭据的跨域请求仍然受第三方 Cookie 策略约束，服务端和前端配置无法绕过用户代理的策略。</p>\n<p>今天至少要按浏览器分别理解：</p>\n<ul>\n<li>Safari/WebKit 很早就默认限制并阻止大量第三方 Cookie，2020 年已经进入完整第三方 Cookie 阻止阶段。</li>\n<li>Firefox 的增强跟踪保护会阻止一部分跟踪类第三方 Cookie。</li>\n<li>Chrome 在 2025 年宣布继续保留用户对第三方 Cookie 的选择，不再推出新的独立提示；但无痕模式默认阻止第三方 Cookie，用户也可以在隐私设置里关闭第三方 Cookie。</li>\n</ul>\n<p>所以，不应再把第三方 Cookie 当成稳定的登录基础设施。对普通业务系统，最稳妥的是尽量避免“前端站点和登录 Cookie 所在接口站点完全不同站”的设计。</p>\n<p>如果确实是在做第三方嵌入组件，比如 iframe 小组件、地图、客服、支付或跨站嵌入应用，可以再评估 Storage Access API、CHIPS/Partitioned Cookie 等方案。但这些方案有明确场景边界，不适合作为普通前后端分离登录的默认解法。</p>\n<h2>方案选择</h2>\n<p>按稳定性排序，我会这样选：</p>\n<ol>\n<li><strong>同源部署</strong>：前端和 API 放在同一个 origin，或者用 Nginx/BFF 把 <code>/api</code> 代理到后端。Cookie 最简单，CORS 问题也最少。</li>\n<li><strong>同站不同源</strong>：例如 <code>www.example.com</code> + <code>api.example.com</code>。需要 CORS 和 <code>credentials: 'include'</code>，但 Cookie 仍在同站语义内。</li>\n<li><strong>不同站但不用 Cookie 做接口身份</strong>：开放平台、跨组织 API、移动端 API 更适合用 OAuth、短期 token、Authorization header 等方式。</li>\n<li><strong>不同站且必须用 Cookie</strong>：只有在明确知道浏览器兼容性、用户设置和嵌入场景的情况下再做，并准备好第三方 Cookie 被禁用时的降级方案。</li>\n</ol>\n<p>反向代理不是“土办法”。对自己控制的 Web 应用来说，把浏览器看到的前端和 API 收敛到同一个站点下，通常比和浏览器隐私策略对抗更稳。</p>\n<h2>安全边界</h2>\n<p>Cookie 会被浏览器自动带上，这也是 CSRF 的基础。CORS 不是 CSRF 防护。一个跨站表单提交或简单请求可以发出去，只是攻击页面不一定能读到响应。</p>\n<p>如果接口使用 Cookie 做登录态，至少要考虑：</p>\n<ul>\n<li>会话 Cookie 使用 <code>HttpOnly; Secure</code>。</li>\n<li>能用 <code>SameSite=Lax</code> 就不要用 <code>SameSite=None</code>。</li>\n<li>对会改变状态的请求校验 CSRF token，或校验 <code>Origin</code>/<code>Sec-Fetch-Site</code> 等请求来源信号。</li>\n<li>不要把 <code>Access-Control-Allow-Origin</code> 无脑反射所有 <code>Origin</code>。</li>\n<li>不要在带凭据的 CORS 响应里使用 <code>Access-Control-Allow-Origin: *</code>。</li>\n</ul>\n<p>Cookie 解决的是身份自动携带，不等于请求就是可信的。</p>\n<h2>调试清单</h2>\n<p>排查 CORS 下 Cookie 收不到时，可以按这个顺序看：</p>\n<ol>\n<li>请求是不是跨源：scheme、host、port 是否完全一致。</li>\n<li>前端是否设置了 <code>fetch(..., { credentials: 'include' })</code> 或 <code>withCredentials = true</code>。</li>\n<li>响应是否有明确的 <code>Access-Control-Allow-Origin</code>，且不是 <code>*</code>。</li>\n<li>响应是否有 <code>Access-Control-Allow-Credentials: true</code>。</li>\n<li>动态 origin 是否加了 <code>Vary: Origin</code>。</li>\n<li><code>Set-Cookie</code> 是否来自接口域名，而不是 response body。</li>\n<li>Cookie 的 <code>Domain</code>、<code>Path</code> 是否覆盖了下一次请求的 URL。</li>\n<li>跨站请求是否设置 <code>SameSite=None; Secure</code>。</li>\n<li>生产环境是否是 HTTPS。</li>\n<li>浏览器是否阻止了第三方 Cookie。</li>\n<li>DevTools 的 Network 请求里是否显示 Cookie 被 blocked，以及 blocked reason。</li>\n<li>Application 面板里 Cookie 所属站点是否符合预期。</li>\n</ol>\n<p>Chrome DevTools 里，Network 面板点开具体请求，看 <code>Cookies</code> 子面板通常比只看 <code>Headers</code> 更清楚。被 SameSite、Secure、Domain、第三方 Cookie 策略拦掉的 Cookie，往往会在这里或 Issues 面板里给出原因。</p>\n<h2>总结</h2>\n<p>CORS 下 Cookie 能不能生效，取决于一整条链路：</p>\n<ul>\n<li>前端要允许带凭据。</li>\n<li>服务端要明确允许对应 origin 携带凭据。</li>\n<li>Cookie 要由目标域名通过 <code>Set-Cookie</code> 设置。</li>\n<li>Cookie 的 Domain、Path、SameSite、Secure 要匹配请求场景。</li>\n<li>浏览器第三方 Cookie 策略不能把它拦掉。</li>\n</ul>\n<p>2017 年那次问题的根因是“前端不能替后端域名写 Cookie”。今天再补一句：即使 Cookie 是后端正确设置的，也要把 SameSite、Secure 和第三方 Cookie 限制一起纳入设计。</p>\n<h2>扩展阅读</h2>\n<ul>\n<li><a href=\"https://developer.mozilla.org/en-US/docs/Web/HTTP/Guides/CORS\">MDN: Cross-Origin Resource Sharing</a></li>\n<li><a href=\"https://developer.mozilla.org/en-US/docs/Web/HTTP/Reference/Headers/Set-Cookie\">MDN: Set-Cookie</a></li>\n<li><a href=\"https://developer.mozilla.org/en-US/docs/Web/HTTP/Guides/Cookies\">MDN: Using HTTP cookies</a></li>\n<li><a href=\"https://developer.mozilla.org/en-US/docs/Web/API/Fetch_API/Using_Fetch#including_credentials\">MDN: Using the Fetch API - Including credentials</a></li>\n<li><a href=\"https://developer.chrome.com/docs/devtools/application/cookies/\">Chrome for Developers: View, add, edit, and delete cookies</a></li>\n<li><a href=\"https://privacysandbox.com/news/privacy-sandbox-next-steps/\">Privacy Sandbox: Next steps for Privacy Sandbox and tracking protections in Chrome</a></li>\n<li><a href=\"https://webkit.org/blog/10218/full-third-party-cookie-blocking-and-more/\">WebKit: Full Third-Party Cookie Blocking and More</a></li>\n<li><a href=\"https://developer.mozilla.org/en-US/docs/Web/Privacy/Privacy_sandbox/Partitioned_cookies\">MDN: Cookies Having Independent Partitioned State</a></li>\n</ul>\n","date_published":"2017-09-02T00:00:00.000Z","date_modified":"2026-05-04T00:00:00.000Z","tags":["CORS","Cookie","SameSite","前端","浏览器"],"language":"zh"},{"id":"https://www.lihuanyu.com/posts/2017/%E4%BD%BF%E7%94%A8Docker%E8%A7%A3%E5%86%B3%E5%BC%80%E5%8F%91%E7%8E%AF%E5%A2%83%E9%97%AE%E9%A2%98/","url":"https://www.lihuanyu.com/posts/2017/%E4%BD%BF%E7%94%A8Docker%E8%A7%A3%E5%86%B3%E5%BC%80%E5%8F%91%E7%8E%AF%E5%A2%83%E9%97%AE%E9%A2%98/","title":"使用 Docker 解决开发环境问题","summary":"已并入《重新认识 Docker：开发环境、Linux 性能开销与 Redis 实战》。","content_html":"<p>本文记录的是一次早期 Docker Compose 开发环境实践：用一个 web 容器运行 Spring Boot，用一个 MySQL 容器提供数据库，让项目可以通过一条命令启动。</p>\n<p>完整讨论见：</p>\n<p><a href=\"/posts/2025/%E9%87%8D%E6%96%B0%E8%AE%A4%E8%AF%86Docker%E7%9A%84%E6%80%A7%E8%83%BD%E5%BC%80%E9%94%80/\">重新认识 Docker：开发环境、Linux 性能开销与 Redis 实战</a></p>\n<p>保留这个页面，是为了让原链接仍然可访问。Docker 用来统一开发环境的价值仍然成立，尤其适合数据库、中间件和后端依赖；但今天更需要同时讨论 Docker Desktop 在 macOS/Windows 上的文件系统成本，以及 Docker Engine 在 Linux 服务器上的实际运行开销。</p>\n","date_published":"2017-08-20T00:00:00.000Z","date_modified":"2026-05-04T00:00:00.000Z","tags":["docker","开发环境"],"language":"zh"},{"id":"https://www.lihuanyu.com/posts/2017/package-lock-json-%E8%AF%91/","url":"https://www.lihuanyu.com/posts/2017/package-lock-json-%E8%AF%91/","title":"package-lock.json[译]","summary":"已并入《前端依赖、lockfile 与可信构建》。","content_html":"<p>本文最初是对 npm <code>package-lock.json</code> 文档的翻译。npm lockfile 的格式和行为后来经历过多次演进，早期翻译已经不适合作为今天理解 npm 依赖管理的主要参考。</p>\n<p>完整说明见：</p>\n<p><a href=\"/posts/2022/%E5%89%8D%E7%AB%AF%E4%BE%9D%E8%B5%96%E4%B8%8E%E4%BF%A1%E4%BB%BB/\">前端依赖、lockfile 与可信构建</a></p>\n<p>保留这个页面，是为了让原链接仍然可访问。今天更值得关注的不是逐字段翻译 lockfile，而是它在工程流程中的位置：提交 lockfile、使用 <code>npm ci</code> 或 frozen install、把依赖变化纳入代码审查，并让 CI 和部署环境安装到同一棵依赖树。</p>\n","date_published":"2017-08-10T00:00:00.000Z","date_modified":"2026-05-04T00:00:00.000Z","tags":["npm","lockfile"],"language":"zh"},{"id":"https://www.lihuanyu.com/posts/2017/webpack%E6%A8%A1%E6%9D%BFmock%E6%95%B0%E6%8D%AE%E7%9A%84%E6%96%B9%E6%B3%95/","url":"https://www.lihuanyu.com/posts/2017/webpack%E6%A8%A1%E6%9D%BFmock%E6%95%B0%E6%8D%AE%E7%9A%84%E6%96%B9%E6%B3%95/","title":"vue-cli webpack模板mock数据的方法","summary":"早期 vue-cli webpack 模板 mock 数据方案的归档页。原方案基于修改 dev-server.js 和 Express 路由，在今天已经不适合作为主要实践参考。","content_html":"<p>本文保留为归档记录。</p>\n<p>这篇文章写于 2017 年，背景是当时的 <code>vue-cli</code> webpack 模板。原方案的核心思路，是直接改开发服务器，在 <code>dev-server.js</code> 里给 Express 增加本地路由，让接口请求返回 <code>mock/</code> 目录下的 JSON 文件。</p>\n<p>这在当时是一个能跑的办法。它解决的是一个很具体的问题：后端接口还没好，前端项目又需要独立跑起来。只要本地能返回约定好的假数据，页面、交互和状态逻辑就可以先往前走。</p>\n<p>但今天已经不适合作为主要实践参考。</p>\n<p>原因也简单：脚手架、构建工具和团队协作方式都变了。直接修改脚手架生成的开发服务器文件，短期省事，长期容易和工具升级、环境差异、接口变更缠在一起。mock 数据也不应该只是一堆散落的 JSON，它最好和接口契约、错误场景、权限状态、分页筛选、网络失败一起管理。</p>\n<p>如果现在重新做，通常会优先考虑几类方案：</p>\n<ol>\n<li>使用框架或构建工具提供的 dev server middleware。</li>\n<li>用 MSW 这类工具在浏览器或 Node 层拦截请求。</li>\n<li>从 OpenAPI、接口文档或后端契约生成 mock。</li>\n<li>对关键业务接口补契约测试，避免前后端各说各话。</li>\n<li>把 mock 场景纳入本地开发和自动化测试，而不是只服务页面临时预览。</li>\n</ol>\n<p>早期方案的价值不在具体代码，而在那个意识：前端工程应该能在后端不完整时独立启动，后来接手的人也应该能快速看到页面跑起来。</p>\n<p>这个意识今天仍然成立，只是实现手段该换了。</p>\n","date_published":"2017-07-09T00:00:00.000Z","date_modified":"2026-05-16T00:00:00.000Z","tags":["前端","mock","归档"],"language":"zh"},{"id":"https://www.lihuanyu.com/posts/2017/%E8%A7%A3%E5%86%B3%E9%97%AE%E9%A2%98%E4%B9%8B%E9%81%93/","url":"https://www.lihuanyu.com/posts/2017/%E8%A7%A3%E5%86%B3%E9%97%AE%E9%A2%98%E4%B9%8B%E9%81%93/","title":"排查问题时，不要太早相信第一假设","summary":"从一次 Vue 组件事件失效的排查经历出发，讨论为什么排查问题时不要太早相信第一假设，以及如何用控制变量、断点、DOM 身份和版本控制把问题一步步缩小。","content_html":"<p>工程师解决问题时，最危险的东西之一，是一个看起来很合理的第一假设。</p>\n<p>它来得很快，解释力很强，还常常带着一点经验的光芒。人一旦相信它，就会开始围着它找证据。找不到，也不一定怀疑假设，反而怀疑自己查得还不够深。</p>\n<p>这时问题就麻烦了。</p>\n<p>排查问题最怕的不是没有方向，而是方向错了还走得很坚决。</p>\n<h2>那次事件失效</h2>\n<p>刚工作不久时，遇到过一个 Vue 组件的问题。</p>\n<p>前一天封好的一个移动端二级菜单组件，在一个页面上已经正常使用。第二天放到另一个页面，突然失效。这个菜单要求能左右滑动，也能点击。</p>\n<p>当时为了统一 iOS 和 Android 的滚动体验，没有直接用原生滚动，而是用了 bscroll 这类滚动库。问题页面本身又有纵向滚动，也用了类似的滚动能力。</p>\n<p>于是第一假设非常自然地冒了出来：</p>\n<p>是不是外层滚动库拦截了事件？</p>\n<p>这个假设听起来很像那么回事。滚动库、移动端、事件冒泡、阻止默认行为、点击穿透，几个词往一起一摆，像极了前端问题。</p>\n<p>于是我在事件上倒腾了一下午。</p>\n<p>查事件冒泡，查捕获，查 <code>preventDefault</code>，查 <code>stopPropagation</code>，查外层容器，查滚动库。越查越像一场泥地行军，脚下全是细节，方向却越来越模糊。</p>\n<p>最后发现，问题根本不在事件流。</p>\n<p>真正的原因很低级：组件里用了 id 选择器。</p>\n<p>上一个页面只有一个组件，所以没出事；新页面里有两个同样的组件。id 本来应该唯一，重复以后选择器拿到的是第一个。偏偏第一个组件因为样式问题处于隐藏状态。</p>\n<p>也就是说，我一直在一个看不见的组件上绑定事件，然后在另一个没有绑定事件的组件上疯狂调试。</p>\n<p>这事说出来有点可笑。</p>\n<p>但排查问题时，很多时间就是这样丢掉的。不是丢在难题上，而是丢在一个过早相信的假设上。</p>\n<h2>第一假设为什么危险</h2>\n<p>第一假设通常不是胡说。</p>\n<p>它往往来自经验。滚动库确实可能处理事件，移动端事件确实复杂，嵌套滚动确实容易出问题。正因为它合理，才更危险。</p>\n<p>完全荒谬的猜测，人反而不会信。最容易误导人的，是那种“八成就是它”的判断。</p>\n<p>一旦心里有了这个判断，后面的动作就会变形。</p>\n<p>你会优先看和它有关的代码，会把新现象解释成它的旁证，会忽略那些不符合它的细节。调试从寻找真相，变成替假设辩护。</p>\n<p>这和写 bug 没什么区别。</p>\n<p>写 bug 是代码相信了错误前提；查 bug 是人相信了错误前提。</p>\n<h2>先确认事实，不急着解释</h2>\n<p>后来再排查问题，我会尽量先做一件事：把问题描述成事实，而不是解释。</p>\n<p>不说：</p>\n<blockquote>\n<p>bscroll 把点击事件拦截了。</p>\n</blockquote>\n<p>而说：</p>\n<blockquote>\n<p>在 A 页面点击菜单有效，在 B 页面点击菜单无效。B 页面里目标 DOM 上是否真的绑定了点击事件，还没有确认。</p>\n</blockquote>\n<p>这两句话差别很大。</p>\n<p>前一句已经下结论，后一句只是记录现象。只要还停留在现象层，就更容易继续问问题：</p>\n<ol>\n<li>事件有没有绑定到预期元素？</li>\n<li>绑定的是不是当前看到的那个组件？</li>\n<li>点击时事件有没有触发？</li>\n<li>如果触发了，执行到哪一步停了？</li>\n<li>如果没触发，是 DOM 不对，时机不对，还是被阻止了？</li>\n</ol>\n<p>这些问题比“是不是滚动库搞鬼”更可靠。</p>\n<p>排查问题不是写侦探小说，不需要一开始就有凶手。先把现场勘清楚，凶手有时会自己站出来。</p>\n<h2>DOM 身份很重要</h2>\n<p>那次问题给我最大的教训，是组件里不要随便用全局 id 选择器。</p>\n<p>id 在 HTML 里本来就应该唯一。可组件的意义，恰恰是可以被多次使用。一旦组件里写死 id，复用时就埋了雷。浏览器和 JavaScript 对重复 id 又很宽容，不会立刻炸给你看，只会在某个页面里悄悄选错元素。</p>\n<p>Vue 里可以用 <code>ref</code>。它能在组件实例里定位元素或子组件，不会像全局 id 那样互相打架。</p>\n<p>更重要的是，要始终确认“我操作的是不是我以为的那个东西”。</p>\n<p>这个问题不只存在于 DOM。</p>\n<p>React 里的 <code>key</code>、Vue 里的组件实例、表单里的字段名、后端里的对象 id、数据库里的唯一键，背后都是同一个问题：身份。如果身份认错了，后面的逻辑再复杂也没用。</p>\n<p>一个请求打到了错误环境，一个事件绑到了隐藏元素，一个状态更新了旧实例，一个缓存命中了错误 key，表现出来都可能像玄学。</p>\n<p>其实不是玄学，是认错人。</p>\n<h2>调试事件，不要只盯代码</h2>\n<p>事件问题很适合用浏览器开发者工具查。</p>\n<p>Chrome DevTools 里可以看元素上绑定的事件，也可以在 Sources 面板里的 Event Listener Breakpoints 对事件打断点。比如勾选 click，再去页面上点击，就能看到到底是哪段代码被执行。</p>\n<p>这比盯着代码猜要快。</p>\n<p>很多时候，人看代码会自动脑补执行路径。浏览器不会。它只告诉你事实：有没有绑定，触发了谁，调用栈是什么，在哪一步停下。</p>\n<p><code>preventDefault</code> 和 <code>stopPropagation</code> 也要分清楚。</p>\n<p><code>preventDefault</code> 是取消默认行为，比如阻止链接跳转、阻止表单提交、阻止复选框默认勾选。它不是用来阻止事件继续传播的。</p>\n<p><code>stopPropagation</code> 才是阻止事件继续冒泡或捕获。</p>\n<p>这两个方法当然都可能影响问题，但不要一上来就乱加。乱加这些方法，就像屋里漏水时先把所有门窗都封死，看起来在处理，实际可能把新问题也埋了进去。</p>\n<h2>控制变量不是口号</h2>\n<p>排查问题时常说控制变量。</p>\n<p>这句话说起来很容易，真正难的是：怎么知道哪些变量已经被控制住了？</p>\n<p>我的经验是，尽量把问题缩小到能被验证的程度。</p>\n<p>比如那次事件问题，可以按顺序做这些检查：</p>\n<ol>\n<li>页面里到底有几个目标组件？</li>\n<li>目标 DOM 是否唯一？</li>\n<li>事件是否绑定到当前可见元素？</li>\n<li>点击时断点是否进入处理函数？</li>\n<li>去掉外层滚动库后，问题是否还存在？</li>\n<li>换成 <code>ref</code> 后，问题是否消失？</li>\n</ol>\n<p>每一步都只回答一个问题。</p>\n<p>如果一口气改三处，问题好了也不知道是哪处好的；问题没好，也不知道哪处判断错了。调试时最忌讳把实验做成一锅粥。</p>\n<p>版本控制也很重要。</p>\n<p>当时我其实有 git，却没有立刻想到可以大胆删改验证。怕把代码改坏，是很多新人都会有的心理。可如果版本控制在，分支在，工作区能恢复，就应该用它换取验证速度。</p>\n<p>大胆实验，小心提交。</p>\n<p>这是 git 给调试带来的底气。</p>\n<h2>有人可问也很重要</h2>\n<p>那次最后能定位到问题，也离不开同事提醒。</p>\n<p>刚入行时，身边有没有靠谱的人能问，差别非常大。很多问题自己闷头查一天，别人看一眼就能指出方向。不是别人比你聪明多少，而是他的经验里已经踩过类似的坑。</p>\n<p>这也是为什么新人选择环境时，不能只看业务酷不酷、技术栈新不新。</p>\n<p>有没有人做 code review？有没有成熟的调试习惯？有没有工程规范？遇到问题时，是有人一起拆，还是所有人都在救火？这些东西比“我们用最新框架”重要得多。</p>\n<p>技术成长不是闭门修仙。</p>\n<p>很多时候，是在一次次问题排查里，学会别人怎么想。</p>\n<h2>最后</h2>\n<p>这个问题本身并不高级。</p>\n<p>重复 id、隐藏组件、事件绑错对象。讲出来甚至有点像低级错误合集。</p>\n<p>但它留下的教训一直有用：</p>\n<p>不要太早相信第一假设。</p>\n<p>先确认事实，再解释原因。先看事件有没有绑定，再讨论事件为什么没触发。先确认操作对象是谁，再研究复杂机制。先做小实验，再下大判断。</p>\n<p>工程师不是靠猜中答案解决问题，而是靠一步步排除错误答案。</p>\n<p>很多 bug 看起来像深山，走进去才发现只是门口的牌子写错了。</p>\n","date_published":"2017-04-08T00:00:00.000Z","date_modified":"2026-05-16T00:00:00.000Z","tags":["前端","调试","Vue"],"language":"zh"},{"id":"https://www.lihuanyu.com/posts/2017/hello-world/","url":"https://www.lihuanyu.com/posts/2017/hello-world/","title":"Hello World","summary":"记录从 WordPress 迁移到 Hexo 的起点，以及当时对静态博客、服务器维护、CI/CD 和个人写作空间的想法。","content_html":"<p>趁着新年把blog转型成简洁的hexo博客了，之前的文章本来想迁移过来，但是读了一下感觉都惨不忍睹，算了，大部分都扔掉吧，重新开始！</p>\n<h3>为什么要换blog系统</h3>\n<p>之前的blog是用wordpress搭建的，运行其实也算良好，按理说没有动力来更换blog系统。</p>\n<p>但是之前的wordpress确实存在一些问题：</p>\n<ul>\n<li>性能问题：wordpress是数据库形式的博客系统，每个页面的数据是存储在数据库中的，用户要看到内容需要PHP连接数据库，查询内容，渲染HTML，给用户看。这种模式，对于一些稍大的、多人的blog系统是合适的，对于我这种单人的、内容不多的，就有些不必要的性能损耗了。以hexo这种直接生成静态HTML的方式更加经济高效。</li>\n<li>安全问题：wordpress虽然市场占有率很高，但毕竟是一套开源的PHP程序，属于漏洞高发区，而一堆静态页面，我是没想到可以怎么黑……</li>\n<li>折腾问题：wordpress想稳定运行还是挺麻烦的，PHP/MYSQL/NGINX什么的都要配套装好，我当时是不会的，所以我用的一键包。显然，不去碰诸如nginx、https、node，是没法学到什么新东西的。</li>\n<li>成本问题：快毕业了，毕业后没有学生优惠，没有大把大把的廉价服务器资源了，就得在一台服务器上折腾我的所有东西，wordpress这种脆弱的blog系统显然很容易被我一不小心折腾崩溃。</li>\n</ul>\n<h3>更换过程</h3>\n<p>虽然有从wordpress迁移的插件，但是迁移后hexo生成就失败了，不知道是原文章里的一些文字刚好碰到了关键字还是什么别的原因，考虑到原来的文章的质量比较参差不齐，最后决定手工更换。（就是技术烂，复制粘贴解决问题算了）</p>\n<h3>其它</h3>\n<ul>\n<li>统计：百度统计</li>\n<li>第三方评论： DISQUS</li>\n<li>自动集成/部署：travis CI</li>\n</ul>\n<h3>自动集成（travis）</h3>\n<p>抽空弄好了CI/CD，用Travis，因为是在自己的服务器上，root的密钥不能给，单开了一个叫blog的账户。</p>\n<p>操作过程可见上一篇文章，有一些注意事项，首先，服务器上开一个新的账户叫blog，然后去blog用户目录下建个.ssh文件夹，注意先切到blog用户，否则root用户建立的文件夹，blog用户无法访问，将导致无法登陆。</p>\n<p>本地这边，两套密钥倒腾了半天，很费劲，很蓝瘦，最后发现用本地系统的新建用户来隔离两套密钥就可以了。</p>\n<h3>先这样咯。Hello Hexo。</h3>\n","date_published":"2017-01-27T00:00:00.000Z","tags":[],"language":"zh"},{"id":"https://www.lihuanyu.com/posts/2016/%E5%86%99%E4%BA%86%E4%B8%AA%E5%89%8D%E7%AB%AF%E6%B8%B2%E6%9F%93%E7%9A%84%E6%95%99%E7%A8%8B/","url":"https://www.lihuanyu.com/posts/2016/%E5%86%99%E4%BA%86%E4%B8%AA%E5%89%8D%E7%AB%AF%E6%B8%B2%E6%9F%93%E7%9A%84%E6%95%99%E7%A8%8B/","title":"写了个前端渲染的教程","summary":"早年前端渲染教程的归档页。原文记录了从后端模板渲染走向 AJAX 与浏览器端渲染时的理解，今天更适合作为前端发展阶段的历史记录阅读。","content_html":"<p>本文保留为归档记录。</p>\n<p>2016 年写这篇文章时，前端渲染对我来说还是一个新鲜概念。那时的主要问题很朴素：页面到底应该在服务器上拼好，还是把数据交给浏览器，由 JavaScript 接管交互和渲染？</p>\n<p>今天再看，这个问题已经不适合用“前端渲染优于后端渲染”来回答。后来几年里，SPA、SSR、SSG、同构框架、边缘渲染、小程序、移动端容器都走过一轮。前端渲染解决了当年的一些痛点，也制造了新的复杂度：首屏性能、SEO、状态同步、构建链路、接口治理、权限和错误边界。</p>\n<p>所以这篇文章更适合作为早期理解的切片，而不是今天的技术建议。</p>\n<p>当时留下来的判断仍然有一点价值：Web 前端并不只是写页面，而是在数据、状态、交互和展示之间建立秩序。只是这件事后来变得更复杂，也更不适合被某一种渲染方式包打天下。</p>\n<p>如果想看后来对前后端边界的重新思考，可以读这篇：</p>\n<p><a href=\"/posts/2026/AI%E6%97%B6%E4%BB%A3%E5%89%8D%E5%90%8E%E7%AB%AF%E5%88%86%E7%A6%BB%E4%B8%8D%E8%AF%A5%E5%86%8D%E6%98%AF%E9%BB%98%E8%AE%A4%E9%80%89%E9%A1%B9/\">AI时代，前后端分离不该再是默认选项</a></p>\n","date_published":"2016-12-18T00:00:00.000Z","date_modified":"2026-05-16T00:00:00.000Z","tags":["前端","归档"],"language":"zh"},{"id":"https://www.lihuanyu.com/posts/2016/%E5%AF%B9%E5%89%8D%E5%90%8E%E7%AB%AF%E5%88%86%E7%A6%BB%E7%9A%84%E6%80%9D%E8%80%83/","url":"https://www.lihuanyu.com/posts/2016/%E5%AF%B9%E5%89%8D%E5%90%8E%E7%AB%AF%E5%88%86%E7%A6%BB%E7%9A%84%E6%80%9D%E8%80%83/","title":"对前后端分离的思考","summary":"结合学校易班轻应用实践，记录从静态页面、Spring Boot 动态网页到前后端分离架构的演进原因和取舍。","content_html":"<p>结合在学校做易班轻应用时候的一些思考，记录下我为什么要做前后端分离的历史/原因/意义/效果。</p>\n<h2>历史背景</h2>\n<p>一开始说要给易班做轻应用的时候是懵逼的，什么叫轻应用。于是分三个阶段循序渐进地来做这个事。</p>\n<h3>静态网页</h3>\n<p>首先，琢磨了一下发现，所谓轻应用，好像其实就是个网页嘛。OK，最简单的就是静态网页。</p>\n<p>做了一些诸如大物实验数据辅助处理系统，就是拿js算一些加减乘除、生活查询，就是根据客户端时间判断水房是否开门等等。</p>\n<p>找个服务器配置下LNMP（linux + nginx + mysql + php 是个服务器环境配置一键脚本），静态页面和资源丢上去，完工。</p>\n<h3>动态网页</h3>\n<p>接下来难度升级，做个带后端服务的真正的应用了。因为是给易班做，目标是吸引用户到易班上，这些轻应用我基本没有考虑过自己做用户表，就是用户信息都是直接从易班开放平台获取。</p>\n<p>这个地方涉及到Oauth2.0协议拿数据，易班开放平台做的这一个呢有一些让我觉得不舒服的限制，比如回调地址和应用地址必须是同一个，精确到完整的url，还有必须是点对点调用，就是不能回调到多个位置。（注意这是一个问题）</p>\n<p>之后技术选型，选的是springboot框架，一个神奇Java的框架，特点是开发速度特别快，让人觉得，这还是Java web框架么？它为很多东西提供缺省配置，省去编写复杂的xml文件的时间。</p>\n<p>部署上也极其方便，内嵌了tomcat，最后build出来的是一个jar包。在一台安装了jdk的电脑上输入java -jar xxxx.jar即可运行，不需要考虑tomcat的配置。（就是这一点让我放弃了PHP的laravel）</p>\n<p>拿springboot写了3、4个应用吧。比如抽奖、查询、签到等，都是这个框架，配上模板引擎thymeleaf做的。</p>\n<p>在完成了3、4个应用后，我发现了这种模式的一些问题，为了解决这些问题：</p>\n<h3>前后端分离</h3>\n<p>这个阶段里，我希望前端和后端独立部署，后端砍掉View层，把Controller层暴露出来，以API形式提供服务。</p>\n<p>前端是一种类似于Client的模式，向后端发起请求。后端吐json格式的数据。前端拿到数据后自己去渲染数据到页面上。（前端渲染，参考 {% post_link 写了个前端渲染的教程 %}）</p>\n<p>截至到离开组织，该架构初步实现，完成了一个demo级别的应用。其中后端以springboot作为框架，运行在服务器的8086还是多少端口来着有点记不清了。</p>\n<h2>原因</h2>\n<h3>前端方面</h3>\n<p>为什么后端的View层要被砍掉？因为后端不会专业的前端技能。</p>\n<p>前端上我希望向工程化、组件化、模块化看齐（好像并没有做到QAQ），要求前端工程要使用一些脚手架/脚本/node工具，进行诸如资源压缩合并混淆的工作，也就是前端其实是单独的工程。需要单独打包。</p>\n<p>那么这样的前端生成的结果如果作为V层放到后端，后端同学需要在一堆乱码中找到需要替换的变量用模板语法改写。当然我们也可以让前端同学学习下thymeleaf直接以模板语法来写。</p>\n<p>但是问题还是存在的。</p>\n<p>这将导致，应用的升级无比麻烦。仅仅只是前端样式上的一个小变化，就需要前端先修改，再打包，给后端，后端打包，再部署。必须要求前后端在一起工作，频繁交流，才可以。但这对学生开发团队太难了，都有课的人。</p>\n<h3>后端方面</h3>\n<p>除了前后端必须联调导致修改不便之外，还有一旦部署，再想修改很难这个问题。因为用户信息没有存自己的表，必须走易班授权，上线的产品修改地址要审核，要本地测试得把回调地址指回本地，改好后又要指回去，又要审核一次。我怎么给用户解释应用不见了这个问题……？卒……</p>\n<p>还有一个问题是，我前面说为什么选springboot框架时说过，是因为这个框架简单，开发简单，上手简单，部署简单。部署简单是有代价的！</p>\n<p>没在服务器上配置tomcat，打包又是jar包，相当于每个应用，都是独立的，跑在独立的tomcat容器里，运行3个应用就等于开了三个tomcat，5个应用就是5个tomcat，tomcat本来就比较重，再这么开下去，服务器很快就受不了。（虽然按微服务架构的思路来说就应该把应用划分得足够细粒度，但是确实穷，又不想优化它的部署方式，以一个简单的方式部署有利于后面做持续集成、持续部署，而且我没有充足的服务器资源）</p>\n<h2>结果</h2>\n<p>前后端分离，把所有的应用都写在一起，整合应用，只用一个tomcat装。前端请求不同的接口获得数据。还有把一些config信息从代码里转移到了yml文件里。因为生产环境和开发环境的配置文件不同，再也不用像以前那样手工反复修改代码了。</p>\n<h3>具体做法</h3>\n<p>springboot的controller用@RestController注解，提供json格式的输出，方便快捷变成一个RESTful微服务后端。</p>\n<p>前端项目直接部署到nginx服务器，阅读后端文档，自己请求API拿数据，展示。前端可以自由选择框架，angular1/2，react，vue，ember等等，反正接口在那。</p>\n<p>遇到的问题就是跨域，用CORS解决了。这里还有个坑，就是CORS这种跨域资源共享一般是结合ajax请求来使用，ajax是默认不带cookie的，通过查资料修改CORS的配置解决掉了。暂时没别的问题了。</p>\n<p>后面看时间如果有空可能拿出一个例子讲一下。</p>\n","date_published":"2016-07-23T00:00:00.000Z","tags":[],"language":"zh"}]}