Repository navigation
飞书网关子进程改由预加载过的 forkserver 派生,内存按机器人数大幅下降 - #119
Merged
Merged
Conversation
- 飞书 SDK 一导入就加载全部开放平台接口(约 1 万个模块),每个机器人子进程单独导入 要 170MB 以上;机器人一多,容器内存贴着上限,启动时同时导入还会冲顶 OOM。 - 新增 preload:forkserver 先导入 child/service 与 asyncpg 方言,关掉 SDK 模块级 事件循环(fork 后会共用 epoll),再 gc.freeze();子进程 run_process 里另建循环。 runtime 镜像实测 7 个机器人:约 330MiB(原约 1250MiB)。 - supervisor 不再导入 SDK:子进程入口 run_child_process 延迟导入 child。 - 子进程改看 multiprocessing 的 supervisor 哨兵判断父进程存活(forkserver 被子进程 持有的管道吊着,父进程号不再变化)。 - spawn 被取消时等线程里的启动做完,把进程挂到 child 上,排空时先停进程再释放租约。 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
问题
飞书网关每个机器人一个子进程。飞书 SDK 的包入口会
from .api import *,导入任何子模块都会加载全部开放平台接口(约 1 万个模块),每个子进程单独导入后 PSS 约 176MiB(Linux runtime 镜像实测)。7 个机器人就把 1.5GiB 的容器上限吃满;重启时所有子进程同时导入,CPU 打满、内存冲顶,网关会陷入 OOMKilled 循环,每次重启所有飞书机器人断线。改动
gateway_feishu/preload.py,forkserver 先导入子进程用到的模块,再gc.freeze(),子进程 fork 后共享这些页。不冻结 GC 时,回收会写脏共享页,几乎省不下来。lark_oapi.ws.client导入时建了一个循环,长连接跑在上面;fork 后各子进程会共用同一个 epoll。preload 里先关掉,child.run_process里每个子进程另建一个。run_child_process延迟导入 child,supervisor 自身仍约 76MB。multiprocessing.parent_process().is_alive()。process.start在线程里跑;停机取消调度时进程仍会起来,spawn现在等启动做完并挂到child.process,排空时照常先停进程再释放租约。ForkedProcess把 multiprocessing 进程包成原来的 asyncio 子进程接口。实测(runtime 镜像,Linux)
gc.freeze()后 fork(本 PR)另在容器里验证:子进程用不可达数据库启动时照常记
feishu_child_failed并以 1 退出;supervisor 被 SIGKILL 后parent_process().is_alive()立刻变 False(父进程号不变)。测试
tests/unit/test_feishu_forkserver.py:退出码回报、terminate、子进程确实来自预加载的 forkserver、每个子进程的 SDK 循环独立、启动中被取消时进程挂到 child 上(撤掉修复时该用例失败)。无数据库迁移,无配置变更。
🤖 Generated with Claude Code