离线安装 K3S 实战:用 TaoToken 统一 Key 打通 Helm、MySQL 与 Longhorn 的本地验证链路
1. 内网离线 K3S 到底难在哪从 Helm 离线包到 Longhorn 卷挂载的完整链路内网无外网环境下部署 K3S难点从来不是装不上而是装完之后 Helm 拉不到 chart、镜像拉不下来、Longhorn 的 iSCSI 依赖缺失、MySQL 作为外部数据存储连不上——每一个环节都能让集群卡在半死状态。我这次要复现的场景是一台联网虚拟机做资源准备多台内网服务器做实际部署K3S 版本 v1.23.17k3s1禁用默认组件coredns、servicelb、traefik、local-storage、metrics-server改用 Helm 单独安装 coredns、metrics-server、traefik、longhorn外部 MySQL 存元数据。这套链路的核心检索词就是K3S 离线安装 Helm 离线包导入 Longhorn 存储卷挂载。适合谁适合需要在隔离网络里跑有状态服务MySQL、Redis 之类的运维和平台工程师尤其是那些被离线环境 Helm 装不上折磨过的人。我踩过的坑主要集中在三块一是 HelmChart CRD 的 bootstrap 机制在离线时不会自动去拉 chart必须提前把 tgz 包塞进对应目录二是 Longhorn 依赖 iscsi-initiator-utils 和 nfs-utils内网机器没装就直接卡在AttachVolume阶段三是多工具调用凭证散落在各个组件的 values.yaml 里改一次要翻五个文件。第三点后面会用 TaoToken 统一 Key 来解决。先说清楚整体架构联网虚拟机上跑一个单节点 K3S用它来拉镜像、导出 chart、下载二进制然后把所有资源打包传到内网服务器内网服务器用外部 MySQL 做 datastore引导 server 和 agent 节点最后通过 HelmChart CRD 或 helm install 把应用装上去。整个过程中凡是需要调用外部 API 的地方比如后续接模型服务、CI 触发部署统一走 TaoToken 的 API 通道Key 只维护一份。2. TaoToken 前置统一 Key 与 API 通道让离线集群鉴权一致离线集群本身不依赖外网但集群里跑的应用、CI 流水线、以及你本地用来调试的 coding agent往往需要调用外部模型或 API。如果每个工具各配一套 Key改起来就是灾难。TaoToken 在这里的角色是统一 Key 管理 统一 API 通道你只需要在 TaoToken 控制台生成一个 Key然后在各个工具里填同一个 Base URL 和 Key 就行。具体来说TaoToken 提供的能力包括模型对话用于调试和验证、Coding Plan长期编码/Agent 场景、API Keys 管理、以及接入文档。对于离线 K3S 场景最实用的两个入口是API Keys 管理页https://taotoken.net/api-keys?utm_sourcetaotoken_aicg_blog_endutm_contentapi_keysutm_campaignrewrite—— 在这里生成和轮换 Key。接入文档https://taotoken.net/doc?utm_sourcetaotoken_aicg_blog_endutm_contentdocutm_campaignrewrite—— 查 Base URL、Model ID、各工具配置方式。如果你要验证模型是否通用模型对话页https://taotoken.net/?utm_sourcetaotoken_aicg_blog_endutm_contentmodel_chatutm_campaignrewrite。长期编码或 Agent 场景走 Coding Planhttps://taotoken.net/coding-plan?utm_sourcetaotoken_aicg_blog_endutm_contentcoding_planutm_campaignrewrite。关键点TaoToken 的 API 地址是https://taotoken.net/api不加 UTM所有工具统一填这个 Base URL。这样你在 K3S 集群里跑的 CI Job、本地的 Claude Code、Cline MCP 配置用的都是同一个 Key 和同一个通道鉴权逻辑一致排查问题时只需要看一个地方。对于离线集群你不需要在每台内网机器上配 Key——只需要在需要调用外部 API 的那个 Pod 或 Job 里注入环境变量。比如你有一个 CI Runner 跑在 K3S 里它需要调模型做代码审查那就在它的 Deployment 里加一个 Secret值从 TaoToken 拿。这样 Key 的轮换只改一个 Secret不用动集群里其他东西。3. 可复制配置HelmChart CRD、MySQL 连接与 Longhorn values 片段这一节给出可以直接复制的配置片段。路径和原文保持一致方便你对照。3.1 K3S server 启动参数含外部 MySQL在联网虚拟机上先跑一个单节点集群做资源准备命令如下curl -fsSL https://rancher-mirror.oss-cn-beijing.aliyuncs.com/k3s/k3s-install.sh | \ INSTALL_K3S_MIRRORcn INSTALL_K3S_VERSIONv1.23.17k3s1 bash -s - server \ --data-dir /data/k3s/var/lib/rancher/k3s \ --cluster-cidr 10.8.0.0/16 \ --service-cidr 10.16.0.0/16 \ --cluster-dns 10.16.0.10 \ --service-node-port-range 1-65535 \ --kube-proxy-arg proxy-modeipvs \ --disable coredns \ --disable servicelb \ --disable traefik \ --disable local-storage \ --disable metrics-server内网 server 节点加上外部 MySQL datastoreINSTALL_K3S_MIRRORcn INSTALL_K3S_VERSIONv1.23.17k3s1 ./install.sh server \ --data-dir /data/k3s/var/lib/rancher/k3s \ --cluster-cidr 10.8.0.0/16 \ --service-cidr 10.16.0.0/16 \ --cluster-dns 10.16.0.10 \ --service-node-port-range 1-65535 \ --kube-proxy-arg proxy-modeipvs \ --disable coredns \ --disable servicelb \ --disable traefik \ --disable local-storage \ --disable metrics-server \ --datastore-endpointmysql://USERNAME:PASSWORDtcp(HOST:3306)/DATABASE注意所有 server 节点的--cluster-cidr、--service-cidr、--cluster-dns必须一致否则节点加入会失败。3.2 HelmChart CRD 配置coredns / metrics-server / traefik / longhorn离线环境下HelmChart CRD 的bootstrap: true不会去外网拉 chart你需要提前把 tgz 包放到/var/lib/rancher/k3s/server/static/charts/目录对应>apiVersion: helm.cattle.io/v1 kind: HelmChart metadata: name: coredns namespace: kube-system labels: app: coredns spec: repo: https://coredns.github.io/helm chart: coredns targetNamespace: kube-system bootstrap: true valuesContent: |- fullnameOverride: coredns serviceType: ClusterIP service: clusterIP: 10.16.0.10 name: coredns servers: - zones: - zone: . port: 53 plugins: - name: errors - name: health configBlock: |- lameduck 5s - name: ready - name: kubernetes parameters: cluster.local in-addr.arpa ip6.arpa configBlock: |- pods insecure fallthrough in-addr.arpa ip6.arpa ttl 30 - name: prometheus parameters: 0.0.0.0:9153 - name: forward parameters: . /etc/resolv.conf - name: cache parameters: 30 - name: loop - name: reload - name: loadbalancemetrics-server 配置apiVersion: helm.cattle.io/v1 kind: HelmChart metadata: name: metrics-server namespace: kube-system labels: app: metrics-server spec: repo: https://charts.bitnami.com/bitnami chart: metrics-server targetNamespace: kube-system bootstrap: true valuesContent: | apiService: create: true extraArgs: - --kubelet-insecure-tls - --kubelet-use-node-status-port - --kubelet-preferred-address-typesInternalIP,ExternalIP,Hostname - --metric-resolution15straefik 配置注意 namespace 是 traefik-systemapiVersion: v1 kind: Namespace metadata: name: traefik-system --- apiVersion: helm.cattle.io/v1 kind: HelmChart metadata: name: traefik namespace: traefik-system labels: app: traefik spec: repo: https://traefik.github.io/charts chart: traefik targetNamespace: traefik-system bootstrap: true valuesContent: |- deployment: kind: Deployment ingressClass: enabled: true isDefaultClass: true providers: kubernetesCRD: enabled: true allowCrossNamespace: true allowExternalNameServices: true allowEmptyServices: true kubernetesIngress: enabled: true allowExternalNameServices: true allowEmptyServices: true publishedService: enabled: true ports: traefik: port: 9000 protocol: TCP expose: false exposedPort: 9000 metrics: port: 9100 protocol: TCP expose: false exposedPort: 9100 web: port: 80 protocol: TCP expose: true exposedPort: 80 nodePort: 30080 websecure: port: 443 protocol: TCP expose: true exposedPort: 443 nodePort: 30443 tls: enabled: true service: type: NodePort securityContext: capabilities: drop: [] add: [ALL] readOnlyRootFilesystem: false podSecurityContext: runAsGroup: 0 runAsNonRoot: false runAsUser: 0longhorn 配置关键defaultDataPath 和副本数apiVersion: v1 kind: Namespace metadata: name: longhorn-system --- apiVersion: helm.cattle.io/v1 kind: HelmChart metadata: name: longhorn namespace: longhorn-system labels: app: longhorn spec: repo: https://charts.longhorn.io chart: longhorn targetNamespace: longhorn-system bootstrap: true valuesContent: |- persistence: defaultClassReplicaCount: 1 csi: attacherReplicaCount: 1 provisionerReplicaCount: 1 resizerReplicaCount: 1 snapshotterReplicaCount: 1 defaultSettings: defaultDataPath: /data/longhorn defaultReplicaCount: 1 deletingConfirmationFlag: true longhornUI: replicas: 1 longhornConversionWebhook: replicas: 1 longhornAdmissionWebhook: replicas: 1 longhornRecoveryBackend: replicas: 1 ingress: enabled: true host: longhorn.example.org3.3 统一 Key 注入TaoToken在需要调用外部 API 的 Pod 里用 Secret 注入 TaoToken 的 KeyapiVersion: v1 kind: Secret metadata: name: taotoken-credentials namespace: default type: Opaque stringData: TAOTOKEN_API_KEY: sk-你的Key TAOTOKEN_BASE_URL: https://taotoken.net/api --- apiVersion: apps/v1 kind: Deployment metadata: name: ci-runner namespace: default spec: replicas: 1 selector: matchLabels: app: ci-runner template: metadata: labels: app: ci-runner spec: containers: - name: runner image: your-ci-runner:latest envFrom: - secretRef: name: taotoken-credentials这样 CI Runner 里所有需要调模型的地方都读TAOTOKEN_API_KEY和TAOTOKEN_BASE_URLKey 轮换只改 Secret。4. 验证请求与成功结果kubectl 命令逐项确认配置写完接下来是验证。离线环境最怕看起来装上了但实际没工作所以每一步都要有明确的成功标志。4.1 确认 K3S 节点和 datastorekubectl get nodes -o wide期望输出所有 server 和 agent 节点状态为ReadyROLES 列显示control-plane,master或worker。确认外部 MySQL datastore 生效cat /data/k3s/var/lib/rancher/k3s/server/token能读到 token 说明 server 正常启动。再检查 MySQL 里是否建了表mysql -h HOST -u USERNAME -p -e USE DATABASE; SHOW TABLES;期望看到kine相关的表说明 K3S 确实在用外部 MySQL 存元数据。4.2 确认 HelmChart 应用状态kubectl get helmchart -A期望输出coredns、metrics-server、traefik、longhorn 四个 chart 的STATUS为deployed。如果卡在pending-install多半是 chart tgz 没放到 static/charts 目录。检查 Podkubectl get pods -n kube-system kubectl get pods -n traefik-system kubectl get pods -n longhorn-systemcoredns 应该是 Runningmetrics-server 应该是 Runningtraefik 应该是 Runninglonghorn 的longhorn-manager、longhorn-driver-deployer、csi-attacher等应该是 Running。4.3 验证 Longhorn 存储卷挂载先确认 Longhorn 的 StorageClass 存在kubectl get sc期望看到longhorn这个 StorageClass并且标注为 default。创建一个 PVC 测试cat EOF | kubectl apply -f - apiVersion: v1 kind: PersistentVolumeClaim metadata: name: test-pvc spec: accessModes: - ReadWriteOnce storageClassName: longhorn resources: requests: storage: 1Gi EOF检查 PVC 状态kubectl get pvc test-pvc期望STATUS为Bound。如果一直是Pending去 longhorn-system 里看longhorn-manager的日志大概率是 iSCSI 依赖没装。再起一个 Pod 挂载这个 PVCcat EOF | kubectl apply -f - apiVersion: v1 kind: Pod metadata: name: test-pod spec: containers: - name: app image: busybox command: [sh, -c, echo hello /data/test.txt sleep 3600] volumeMounts: - name: data mountPath: /data volumes: - name: data persistentVolumeClaim: claimName: test-pvc EOF验证写入kubectl exec test-pod -- cat /data/test.txt期望输出hello说明 Longhorn 卷挂载成功。4.4 验证 MySQL 有状态服务如果你要在 K3S 里跑 MySQL StatefulSet用 Longhorn 做存储apiVersion: apps/v1 kind: StatefulSet metadata: name: mysql spec: serviceName: mysql replicas: 1 selector: matchLabels: app: mysql template: metadata: labels: app: mysql spec: containers: - name: mysql image: mysql:8.0 env: - name: MYSQL_ROOT_PASSWORD value: yourpassword ports: - containerPort: 3306 volumeMounts: - name: data mountPath: /var/lib/mysql volumeClaimTemplates: - metadata: name: data spec: accessModes: [ReadWriteOnce] storageClassName: longhorn resources: requests: storage: 10Gi验证kubectl get pods -l appmysql kubectl exec mysql-0 -- mysql -uroot -pyourpassword -e SHOW DATABASES;期望看到information_schema、mysql、performance_schema等库说明 MySQL 在 Longhorn 卷上正常读写。4.5 验证 TaoToken 通道在 CI Runner Pod 里测试kubectl exec -it ci-runner-xxx -- sh curl -s https://taotoken.net/api/v1/models \ -H Authorization: Bearer $TAOTOKEN_API_KEY期望返回模型列表 JSON。如果返回 401检查 Key 是否正确如果返回连接超时检查集群是否能出网离线集群需要单独配置出口。5. 本篇常见错排查401、local proxy failed、reading choices、OAuth离线 K3S 部署过程中报错集中在几个地方。下面按真实报错对照排查。5.1 401 Unauthorized现象调用 TaoToken API 返回{error:{message:Invalid API key,type:invalid_request_error}}。原因Key 填错、Key 被轮换、或者 Secret 没正确挂载。排查kubectl get secret taotoken-credentials -o jsonpath{.data.TAOTOKEN_API_KEY} | base64 -d确认输出的 Key 和 TaoToken 控制台里的一致。如果不一致重新生成 Secretkubectl create secret generic taotoken-credentials \ --from-literalTAOTOKEN_API_KEYsk-你的新Key \ --from-literalTAOTOKEN_BASE_URLhttps://taotoken.net/api \ --dry-runclient -o yaml | kubectl apply -f -然后重启 Pod。5.2 local proxy failed现象CI Runner 或本地 coding agent 报local proxy failed或connection refused。原因Base URL 填错或者本地代理配置指向了不存在的地址。排查确认 Base URL 是https://taotoken.net/api不要多加/v1或漏掉/api。如果你用的是 Claude Code 或 Cline检查它们的配置文件里 Base URL 字段。对于 Claude Code配置在~/.claude/settings.json{ env: { ANTHROPIC_BASE_URL: https://taotoken.net/api, ANTHROPIC_API_KEY: sk-你的Key } }对于 Cline MCP配置在cline_mcp_settings.json{ mcpServers: { taotoken: { command: npx, args: [-y, taotoken/mcp-server], env: { TAOTOKEN_API_KEY: sk-你的Key, TAOTOKEN_BASE_URL: https://taotoken.net/api } } } }对于 Codex配置在~/.codex/auth.json{ api_key: sk-你的Key, base_url: https://taotoken.net/api }三件套必须齐全Base URL Key Model ID。Model ID 在 TaoToken 接入文档里查。5.3 reading choices 报错现象调用模型返回error reading choices或unexpected end of JSON input。原因请求体格式不对或者 Model ID 填错。排查确认请求体是标准 OpenAI 格式curl -s https://taotoken.net/api/v1/chat/completions \ -H Authorization: Bearer $TAOTOKEN_API_KEY \ -H Content-Type: application/json \ -d { model: claude-3-5-sonnet, messages: [{role: user, content: hello}] }如果 Model ID 不存在会返回model not found。去 TaoToken 模型对话页确认可用模型列表。5.4 OAuth 相关报错现象Claude Code 或 Codex 报OAuth token expired或authentication failed。原因工具默认走 OAuth 流程但离线环境或统一 Key 场景下应该走 API Key。排查Claude Code 里执行claude logout然后重新配置 API Key 模式。Codex 里检查auth.json是否同时存在 OAuth 和 API Key 字段如果有冲突删掉 OAuth 相关字段只保留api_key和base_url。5.5 Longhorn 卷一直 Pending现象PVC 创建后一直Pendingkubectl describe pvc显示waiting for first consumer或no nodes available。原因iSCSI 依赖没装或者节点上没有/data/longhorn目录。排查# 检查 iSCSI systemctl status iscsid # 检查目录 ls -la /data/longhorn如果 iscsid 没启动systemctl enable iscsid --now如果目录不存在mkdir -p /data/longhorn然后重启 longhorn-manager Pod。5.6 HelmChart 卡在 pending-install现象kubectl get helmchart -A显示pending-installPod 没起来。原因chart tgz 没放到/var/lib/rancher/k3s/server/static/charts/。排查ls /data/k3s/var/lib/rancher/k3s/server/static/charts/确认 coredns、metrics-server、traefik、longhorn 的 tgz 包都在。如果没有从联网虚拟机复制过来scp coredns-1.26.0.tgz root内网IP:/data/k3s/var/lib/rancher/k3s/server/static/charts/然后删除失败的 HelmChart 重新 apply。6. 把 Key 管起来离线集群里多工具鉴权的统一入口离线 K3S 集群跑起来之后真正麻烦的是后续维护。集群里可能有 CI Runner、监控 Agent、日志采集器本地还有 Claude Code、Cline、Codex 这些 coding 工具每个都要调外部 API。如果每个工具配一套 Key轮换一次要改十几个地方还容易漏。TaoToken 的价值就在这里一个 Key一个 Base URL所有工具统一填。你可以在 TaoToken 控制台生成 Key然后在 API Keys 页面管理https://taotoken.net/api-keys?utm_sourcetaotoken_aicg_blog_endutm_contentapi_keysutm_campaignrewrite。接入文档在https://taotoken.net/doc?utm_sourcetaotoken_aicg_blog_endutm_contentdocutm_campaignrewrite里面有各工具的详细配置方式。对于长期编码和 Agent 场景Coding Plan 更划算https://taotoken.net/coding-plan?utm_sourcetaotoken_aicg_blog_endutm_contentcoding_planutm_campaignrewrite。如果你只是想验证模型通不通用模型对话页https://taotoken.net/?utm_sourcetaotoken_aicg_blog_endutm_contentmodel_chatutm_campaignrewrite。回到离线 K3S 场景我的建议是把 TaoToken 的 Key 存成 K3S Secret所有需要调外部 API 的 Pod 通过envFrom引用。这样 Key 轮换只改一个 Secret重启相关 Pod 即可。本地 coding 工具则各自读环境变量TAOTOKEN_API_KEY和TAOTOKEN_BASE_URL保持和集群内一致。最后一步验证在集群里跑一个 Job调 TaoToken 的模型接口确认返回正常kubectl run taotoken-test --rm -it --restartNever \ --imagecurlimages/curl \ --envTAOTOKEN_API_KEYsk-你的Key \ -- curl -s https://taotoken.net/api/v1/models \ -H Authorization: Bearer $TAOTOKEN_API_KEY看到模型列表 JSON 返回说明离线集群到 TaoToken 的通道打通了。至此K3S 离线安装、Helm 离线包导入、MySQL 有状态服务、Longhorn 卷挂载、以及统一 Key 管理这条链路全部验证完毕。
上一篇/下一篇内容由系统自动关联
返回资讯列表 →