Margrop
Articles243
Tags633
Categories6

Categories

1password 24GB VRAM 2K 3.6 Flash 30B dense 4-bit 量化 6-DoF SLAM AC ACP AI Agent AI Coding Assistant AI Tech AI tutor AI 安全 AI 应用 AI 日记 AI编程助手 ALTK-Evolve AMIE AP API API 定价 API 降价 ARC-AGI-3 ASR ATEM chat 模板 Agent Agent Harness Agent Memory Agent 入侵 Agent 工程 Agent 架构 Agent 检索 Agent 沙箱 Agent 系统 Agent 路由 Agentic AI Agentic tools Ai2 Alertmanager AllenAI Android 17 Antigravity AppDaemon AppWorld Aqara Astra Attention Baseten Benchmark CC-Switch CI/CD CLI Tools CLI工具 CPU 推理 Cache Hit Rate Caddy ChatGPT Claude Code Claude Sonnet ClawLoader Code Interpreter Codex ComfyUI Computer Use Cookie 认证 Cosmos-H-Dreams Cost Optimization Cron DFIR DSpark Date DeepSeek DeepSeek V4 Flash Diagrams.net Diary Diffusers Diffusion Docker Efficiency Tools Embedding English FSDP2 Fable 5 Fireworks AI FlashAttention FlashDreams GGUF GLM 5.2 GLM-5.2 GPT-4.1 GPT-5.6 GPT-Live GPT-Red GPU 加速 GPU 性能分析 Gateway Gemini Gemini 3.5 Flash Gemini API Gemini CLI Gemini Omni Flash Gemma 4 12B Gemma Translator Gemma4 GitHub Actions Google Google AI Google Research Google Sheets Grabette Gripette HA HADashboard HF Security Incident Hailuo Hermes Hexo HomeAssistant Hugging Face IBM Research Inference Providers Isaac Lab Java KV cache Kimi K3 Kubernetes LFM2.5 LFM2.5-VL LLM Router LVM‑Thin Late Interaction LeRobot Linux Liquid AI LiquidAI Live Translate LoRA Luna MCP MTP MacOS Magpie TTS Managed Agents Meta Microsoft 365 Copilot MiniMax Mistral Shieldstral Model Routing MuJoCo Warp Multi-Agent Multi-Vector Muse Glimmer MySQL NAS NIM NVIDIA NeMo Automodel Nemotron 3 Embed Newton Nginx Node.js Nunchaku OCR OOM OlmoEarth On-device AI Open Source OpenAI OpenAI 兼容 OpenClaw OpenCode OpenResty OpenWrt PII 检测 Physical AI Pollen Robotics Portainer PostgreSQL ProcessOn Project Astra Prometheus Prompt Caching Prompt Injection Proxmox VE PyTorch Qwen3-VL Qwen3.6 RAG RPC RTEB Real-Time Inference Red Teaming Responses API SNAP SOCKS5 SPED SVDQuant Scientific Computing Self-Forcing Distillation Sentence Transformers Session Sheets canvas Shell Sol Storage Buckets Strands Agents Subagent Surgical Robotics TTS Terra Think button TimeMachine TutorMoments UML Uptime Kuma V4-Pro VPS VoiceEQ WARP WebRTC WebSocket Windows World Foundation Model agent agentic aligenie aliyun annotation aop autofs backup bash bitwarden boot brew browser budget control centos cert certbot charles chat chrome classloader client clone closures cloudflare command commit commoditization container crontab cyber capability demo dependency deploy developer devtools dll dns docker domain download drafter draw drawio dsm dump dylib environment hooks exception fail2ban feign firewall-cmd flow free tier frontier hosted model frp frpc frps fuckgfw full-duplex function gfw git github gperftools gridea grub guardrail lockout gvt-g hacs havcs heap hello hexo hibernate hidpi hoisting homeassistant hosts html htmlparser https huggingface_hub iKuai iMessage image img img2kvm immortalwrt import index inference cost install intel io ios ip iptables iso java javascript jni jnilib jpa js json jsonb jupter jupyterlab jvm k8s kernel key kvm lastpass launchctl learning letsencrypt linux llama.cpp low-code lvm mac mariadb markdown maven md5 microcode mirror modules monitor mount mstsc multimodal mysql n5105 network nfs node node-red nodejs nohup notepad++ npm nssm ntp oop open weights openfeign openssl os ovz packet capture pdf pem perf pip plugin png powerbutton print pro productive struggle pve pvekclean python qcow2 qemu qemu-guest-agent rar reasoning control reasoning slider reboot reflog remote remote desktop renew repo resize retina router runtime safari sata scaffolding scheduled triggers scipy-notebook scoping scp self-play server serverless inference silent test simulated student so speculative decoding spk spring springboot springfox ssh ssl stash string support svg svn swagger sync synology systemctl systemd template terminal txt ubuntu ui undertow unlocker upgrade vLLM vhd vim vm vmdk web windows with worker xml yum zai-org/GLM-5.2 zip 上下文压缩 上下文工程 交换机 人才争夺 代理 企业 AI 优化 低延迟 供应链 健康检查 光猫 免费层 内存 内存优化 内网渗透 分布式推理 分布式训练 医疗 AI 升级 卫星影像 反向代理 反诉 向量检索 启动 告警 告警优化 地球观测 地理空间推理 复盘评测 夏令时 多 token 预测 多智能体 多模态 多模态 Agent 多语言 大厂人才战 大模型评测 天猫精灵 安全 安全事件 安装 定时任务 实时语音 客户端 SDK 容器 导入 小米 屏幕理解 工具审计 工具调用 工具调用拦截 工程团队 工程实践 工程笔记 常用软件 应用市场 延迟优化 开权重 开源权重 开源模型 异常 异步任务 异步委派 微信 心跳 性能优化 成本优化 成本控制 扩散模型 技术 抓包 按 provider 优先级 排查 推理加速 推理速度 推理预算 描述文件 提示词敏感性 故障排查 效率工具 教育数据开源 教育评测 数据工作流 数据流 数据集偏差 文本编码器 旁路由 日志分析 日记 时区 显卡虚拟化 智能家居 智能音箱 服务管理 本地 agent 机器人仿真 机器人学习 机器人数据采集 架构 模块 模型推理 模型评测 模型路由 残存访问 流式推理 流程 流程图 浏览器 漫游 火绒 电信 画图 监控 监控系统 监管 磁盘 稀疏注意力 立体声 端侧 AI 端侧推理 端口 端口冲突 端口扫描 续期 网关 网络 网络风暴 群晖 脚本 脚本优化 腾讯 自动化 自动恢复 自动攻击 自部署 苹果 虚拟机 视觉语言模型 视频生成 视频问诊 认证 证书 评测 评测基准 评测方法学 诉讼 语音 AI 语音 Agent 语音识别 超时 路由 路由器 软件管家 软路由 运维 运维监控 连接保活 连接问题 通信机制 通知 邮件漏发 部署 配置 量化 钉钉 镜像 镜像源 长上下文 长连接 门窗传感器 问题排查 防火墙 阿里云 阿里源 集客 飞书

Hitokoto

Archive

VPS健康检查完全指南:从容器到Gateway的全面检测

VPS健康检查完全指南:从容器到Gateway的全面检测

前言

对于运维工程师来说,健康检查是一项基础但非常重要的工作。无论是个人VPS还是生产环境服务器,定期的健康检查能够帮助我们及时发现问题,避免小问题演变成大故障。本文将分享一套完整的VPS健康检查方案,涵盖容器、Gateway、网络和存储等多个维度。

为什么需要健康检查

在日常运维中,我们经常会遇到以下场景:

  • 刚部署好的服务,运行一段时间后突然罢工
  • 磁盘空间不知不觉被占满,导致服务崩溃
  • 网络连接突然中断,但浑然不觉
  • 内存泄漏导致系统越来越慢

这些问题如果能够在早期发现,就能避免很多不必要的麻烦。健康检查就是发现这些问题的第一道防线。

健康检查维度

1. 容器状态检查

对于使用Docker部署的服务,首先需要确认容器是否正常运行。

1
2
3
4
5
6
7
8
9
10
11
# 查看所有容器状态
docker ps -a

# 只看运行中的容器
docker ps

# 查看特定容器的详细信息
docker inspect 容器名

# 查看容器日志(最近100行)
docker logs --tail 100 容器名

关键指标:

  • 容器是否处于Running状态
  • 重启次数是否异常(如果频繁重启,说明有问题)
  • 内存和CPU使用是否正常

常见容器问题:

问题现象 可能原因 解决方案
容器Exited 进程崩溃/配置错误 查看日志排查
容器Restarting 启动脚本错误/资源不足 检查启动脚本和资源
容器OOM 内存不足 调整内存限制

2. Gateway服务检查

对于OpenClaw这类需要长连接的服务,Gateway的状态至关重要。

1
2
3
4
5
6
7
8
9
10
11
# 检查Gateway进程状态
ps aux | grep openclaw

# 检查Gateway端口是否监听
netstat -tlnp | grep 18789

# 检查Gateway健康状态
curl http://localhost:18789/health

# 查看Gateway日志
tail -f /tmp/openclaw-gateway.log

健康状态响应示例:

1
{"ok":true,"status":"live","version":"2026.3.1"}

如果返回ok为true,说明Gateway运行正常;如果返回false或无响应,说明可能存在问题。

3. 消息通道连接检查

不同的消息通道有不同的检查方式:

钉钉连接检查

1
2
3
4
5
# 查看钉钉连接日志
grep -i "dingtalk" /tmp/openclaw-gateway.log

# 检查WebSocket连接状态
curl http://localhost:18789/api/channels

飞书连接检查

1
2
# 查看飞书连接日志
grep -i "feishu" /tmp/openclaw-gateway.log

Discord连接检查

1
2
# 查看Discord连接日志
grep -i "discord" /tmp/openclaw-gateway.log

4. 磁盘空间检查

磁盘空间不足是一个常见但严重的问题。

1
2
3
4
5
6
7
8
9
# 查看磁盘使用情况
df -h

# 查看特定目录的大小
du -sh /var/log/*
du -sh /root/*

# 查看大文件
find / -type f -size +100M -exec ls -lh {} \;

推荐做法:

  • 系统盘使用率不超过80%
  • 数据盘使用率不超过90%
  • 定期清理日志文件
  • 设置告警阈值(建议80%告警,90%严重)

5. 内存和CPU检查

1
2
3
4
5
6
7
8
# 查看内存使用情况
free -h

# 查看CPU使用情况
top -bn1 | head -20

# 查看系统负载
uptime

6. 网络连通性检查

1
2
3
4
5
6
7
8
9
10
11
# 测试基础网络连通性
ping -c 4 8.8.8.8

# 测试DNS解析
nslookup google.com

# 测试特定端口连通性
telnet 目标IP 端口

# 查看网络连接状态
ss -tunap

自动化健康检查方案

手动检查毕竟效率低下,建议搭建自动化健康检查系统。

1. 使用定时任务(Cron)

1
2
# 每天早上9点执行健康检查
0 9 * * * /opt/openclaw/scripts/health-check.sh >> /var/log/health-check.log 2>&1

健康检查脚本示例:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
#!/bin/bash

# 健康检查脚本

# 1. 检查Docker容器状态
echo "=== 容器状态 ==="
docker ps --format "table {{.Names}}\t{{.Status}}\t{{.Ports}}"

# 2. 检查Gateway状态
echo -e "\n=== Gateway状态 ==="
curl -s http://localhost:18789/health

# 3. 检查磁盘空间
echo -e "\n=== 磁盘使用 ==="
df -h | grep -E "^/dev"

# 4. 检查内存
echo -e "\n=== 内存使用 ==="
free -h

# 5. 检查系统负载
echo -e "\n=== 系统负载 ==="
uptime

2. 使用监控系统(Prometheus + Grafana)

对于更复杂的场景,建议使用专业的监控系统:

  • Prometheus: 收集指标数据
  • Grafana: 可视化展示
  • Alertmanager: 告警通知

关键指标:

指标 告警阈值 严重阈值
磁盘使用率 >80% >90%
内存使用率 >80% >90%
CPU使用率 >70% >90%
Gateway进程 Down -
容器状态 Exited -

3. 使用OpenClaw内置健康检查

OpenClaw Gateway提供了健康检查接口,可以集成到自己的监控系统中:

1
2
3
4
5
# 检查Gateway基本状态
curl http://<Gateway-IP>:18789/health

# 检查通道连接状态
curl http://<Gateway-IP>:18789/api/channels

常见问题解答

Q:健康检查频率应该怎么设置?

A:一般建议:

  • 高优先级服务:每5分钟检查一次
  • 中优先级服务:每15分钟检查一次
  • 低优先级服务:每小时检查一次

Q:告警太多怎么办?

A:可以设置告警聚合和静默期:

  • 相同告警在24小时内不重复通知
  • 非工作时间设置静默期
  • 根据告警级别设置不同的通知方式

Q:容器经常自动重启怎么办?

A:按以下步骤排查:

  1. 查看容器日志:docker logs 容器名
  2. 检查容器健康状态:docker inspect 容器名
  3. 检查资源限制是否合理
  4. 检查应用本身是否有bug

Q:磁盘空间总是很快占满怎么办?

A:建议:

  1. 定期清理日志(使用logrotate)
  2. 定期清理临时文件
  3. 监控大文件目录
  4. 考虑扩容或清理历史数据

最佳实践总结

  1. 建立检查清单: 每次检查都按照清单来,避免遗漏

  2. 自动化优先: 尽量用脚本代替手动检查

  3. 设置合理告警: 不要太多也不要太少

  4. 记录历史数据: 方便对比和分析趋势

  5. 定期复盘: 分析告警原因,优化检查策略

结语

健康检查是运维工作的基础,一个好的健康检查体系能够大大降低故障发生的概率,提高系统的稳定性。本文介绍的方法适用于大多数VPS场景,你可以根据自己的实际情况进行调整和优化。

如果你也有好的健康检查经验,欢迎在评论区分享!


作者:小六,一个在上海努力搬砖的程序员

本文阅读量 --
Author:Margrop
Link:https://blog.margrop.com/post/2026-03-13-vps-health-check-guide/
版权声明:本文采用 CC BY-NC-SA 3.0 CN 协议进行许可