Margrop
Articles243
Tags633
Categories6

Categories

1password 24GB VRAM 2K 3.6 Flash 30B dense 4-bit 量化 6-DoF SLAM AC ACP AI Agent AI Coding Assistant AI Tech AI tutor AI 安全 AI 应用 AI 日记 AI编程助手 ALTK-Evolve AMIE AP API API 定价 API 降价 ARC-AGI-3 ASR ATEM chat 模板 Agent Agent Harness Agent Memory Agent 入侵 Agent 工程 Agent 架构 Agent 检索 Agent 沙箱 Agent 系统 Agent 路由 Agentic AI Agentic tools Ai2 Alertmanager AllenAI Android 17 Antigravity AppDaemon AppWorld Aqara Astra Attention Baseten Benchmark CC-Switch CI/CD CLI Tools CLI工具 CPU 推理 Cache Hit Rate Caddy ChatGPT Claude Code Claude Sonnet ClawLoader Code Interpreter Codex ComfyUI Computer Use Cookie 认证 Cosmos-H-Dreams Cost Optimization Cron DFIR DSpark Date DeepSeek DeepSeek V4 Flash Diagrams.net Diary Diffusers Diffusion Docker Efficiency Tools Embedding English FSDP2 Fable 5 Fireworks AI FlashAttention FlashDreams GGUF GLM 5.2 GLM-5.2 GPT-4.1 GPT-5.6 GPT-Live GPT-Red GPU 加速 GPU 性能分析 Gateway Gemini Gemini 3.5 Flash Gemini API Gemini CLI Gemini Omni Flash Gemma 4 12B Gemma Translator Gemma4 GitHub Actions Google Google AI Google Research Google Sheets Grabette Gripette HA HADashboard HF Security Incident Hailuo Hermes Hexo HomeAssistant Hugging Face IBM Research Inference Providers Isaac Lab Java KV cache Kimi K3 Kubernetes LFM2.5 LFM2.5-VL LLM Router LVM‑Thin Late Interaction LeRobot Linux Liquid AI LiquidAI Live Translate LoRA Luna MCP MTP MacOS Magpie TTS Managed Agents Meta Microsoft 365 Copilot MiniMax Mistral Shieldstral Model Routing MuJoCo Warp Multi-Agent Multi-Vector Muse Glimmer MySQL NAS NIM NVIDIA NeMo Automodel Nemotron 3 Embed Newton Nginx Node.js Nunchaku OCR OOM OlmoEarth On-device AI Open Source OpenAI OpenAI 兼容 OpenClaw OpenCode OpenResty OpenWrt PII 检测 Physical AI Pollen Robotics Portainer PostgreSQL ProcessOn Project Astra Prometheus Prompt Caching Prompt Injection Proxmox VE PyTorch Qwen3-VL Qwen3.6 RAG RPC RTEB Real-Time Inference Red Teaming Responses API SNAP SOCKS5 SPED SVDQuant Scientific Computing Self-Forcing Distillation Sentence Transformers Session Sheets canvas Shell Sol Storage Buckets Strands Agents Subagent Surgical Robotics TTS Terra Think button TimeMachine TutorMoments UML Uptime Kuma V4-Pro VPS VoiceEQ WARP WebRTC WebSocket Windows World Foundation Model agent agentic aligenie aliyun annotation aop autofs backup bash bitwarden boot brew browser budget control centos cert certbot charles chat chrome classloader client clone closures cloudflare command commit commoditization container crontab cyber capability demo dependency deploy developer devtools dll dns docker domain download drafter draw drawio dsm dump dylib environment hooks exception fail2ban feign firewall-cmd flow free tier frontier hosted model frp frpc frps fuckgfw full-duplex function gfw git github gperftools gridea grub guardrail lockout gvt-g hacs havcs heap hello hexo hibernate hidpi hoisting homeassistant hosts html htmlparser https huggingface_hub iKuai iMessage image img img2kvm immortalwrt import index inference cost install intel io ios ip iptables iso java javascript jni jnilib jpa js json jsonb jupter jupyterlab jvm k8s kernel key kvm lastpass launchctl learning letsencrypt linux llama.cpp low-code lvm mac mariadb markdown maven md5 microcode mirror modules monitor mount mstsc multimodal mysql n5105 network nfs node node-red nodejs nohup notepad++ npm nssm ntp oop open weights openfeign openssl os ovz packet capture pdf pem perf pip plugin png powerbutton print pro productive struggle pve pvekclean python qcow2 qemu qemu-guest-agent rar reasoning control reasoning slider reboot reflog remote remote desktop renew repo resize retina router runtime safari sata scaffolding scheduled triggers scipy-notebook scoping scp self-play server serverless inference silent test simulated student so speculative decoding spk spring springboot springfox ssh ssl stash string support svg svn swagger sync synology systemctl systemd template terminal txt ubuntu ui undertow unlocker upgrade vLLM vhd vim vm vmdk web windows with worker xml yum zai-org/GLM-5.2 zip 上下文压缩 上下文工程 交换机 人才争夺 代理 企业 AI 优化 低延迟 供应链 健康检查 光猫 免费层 内存 内存优化 内网渗透 分布式推理 分布式训练 医疗 AI 升级 卫星影像 反向代理 反诉 向量检索 启动 告警 告警优化 地球观测 地理空间推理 复盘评测 夏令时 多 token 预测 多智能体 多模态 多模态 Agent 多语言 大厂人才战 大模型评测 天猫精灵 安全 安全事件 安装 定时任务 实时语音 客户端 SDK 容器 导入 小米 屏幕理解 工具审计 工具调用 工具调用拦截 工程团队 工程实践 工程笔记 常用软件 应用市场 延迟优化 开权重 开源权重 开源模型 异常 异步任务 异步委派 微信 心跳 性能优化 成本优化 成本控制 扩散模型 技术 抓包 按 provider 优先级 排查 推理加速 推理速度 推理预算 描述文件 提示词敏感性 故障排查 效率工具 教育数据开源 教育评测 数据工作流 数据流 数据集偏差 文本编码器 旁路由 日志分析 日记 时区 显卡虚拟化 智能家居 智能音箱 服务管理 本地 agent 机器人仿真 机器人学习 机器人数据采集 架构 模块 模型推理 模型评测 模型路由 残存访问 流式推理 流程 流程图 浏览器 漫游 火绒 电信 画图 监控 监控系统 监管 磁盘 稀疏注意力 立体声 端侧 AI 端侧推理 端口 端口冲突 端口扫描 续期 网关 网络 网络风暴 群晖 脚本 脚本优化 腾讯 自动化 自动恢复 自动攻击 自部署 苹果 虚拟机 视觉语言模型 视频生成 视频问诊 认证 证书 评测 评测基准 评测方法学 诉讼 语音 AI 语音 Agent 语音识别 超时 路由 路由器 软件管家 软路由 运维 运维监控 连接保活 连接问题 通信机制 通知 邮件漏发 部署 配置 量化 钉钉 镜像 镜像源 长上下文 长连接 门窗传感器 问题排查 防火墙 阿里云 阿里源 集客 飞书

Hitokoto

Archive

Prometheus 监控显示服务 UP 但实际不可用的排查与修复实践

Prometheus 监控显示服务 UP 但实际不可用的排查与修复实践

前言

在使用 Prometheus 进行服务监控时,你是否遇到过这种情况:监控面板显示服务状态为 “UP”,但实际调用时却返回错误?本文将详细记录一次典型的”监控与实际不符”问题的完整排查过程,并提供实用的修复方案。

问题现象

监控面板状态

在 Prometheus 监控面板中,目标服务的健康检查状态显示为 “UP”,服务响应时间也在正常范围内:

  • app-gateway: UP
  • bot-svr: UP
  • openapi-svr: UP
  • robot-svr: UP
  • robot-task: UP

实际服务状态

但当通过实际业务接口进行测试时,发现:

  • 部分接口返回 503 Service Unavailable
  • 某些环境的 API 调用超时
  • 日志中存在大量错误信息

问题分析

1. 健康检查机制分析

Prometheus 的监控通常依赖于服务暴露的 /actuator/prometheus 端点。这种监控方式存在以下局限性:

监控指标 检查内容 局限性
服务进程状态 进程是否存活 无法检测业务逻辑错误
端口监听状态 端口是否开放 无法检测服务是否真正响应
健康检查端点 /health 是否返回 200 可能只是应用层健康,不代表业务健康

2. 根因定位

经过深入排查,发现以下问题:

问题一:服务实例部分宕机
在某个测试环境中,部分服务实例已经宕机,但负载均衡器的健康检查机制将请求路由到了仅存的其他实例。因为还有实例在响应,所以 Prometheus 的健康检查仍然显示 “UP”。

问题二:HTTP 503 错误的监控盲区
服务返回 503 错误时,HTTP 响应码仍然是 200(因为是健康检查端点返回的),所以监控系统认为服务正常。

问题三:跨区域网络延迟
部分欧洲(prod-xx)和美国(prod-yy)区域的服务存在间歇性超时问题,表现为 “context deadline exceeded”,这是典型的跨境网络延迟问题。

3. 影响范围

根据 Prometheus 监控数据,本次问题影响以下环境:

环境 受影响服务 状态
dev-aliyun bot-svr, openapi-svr, robot-svr, robot-task-trcr DOWN
prod-xx bot-svr, openapi-svr, robot-svr, sch-tsk DOWN
prod-yy robot-rdr DOWN

解决方案

1. 修复服务实例

对于宕机的服务实例,执行以下操作:

1
2
3
4
5
6
7
8
9
10
11
# SSH 到目标服务器
ssh root@<target-server>

# 检查服务状态
systemctl status <service-name>

# 重启服务
systemctl restart <service-name>

# 验证服务是否正常
curl -f http://localhost:8080/health

2. 优化监控配置

方案一:增加业务健康检查

在服务的健康检查端点中添加更全面的检查:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
# application.yml
management:
endpoints:
web:
exposure:
include: health,info,prometheus
endpoint:
health:
show-details: always
probes:
enabled: true
health:
livenessState:
enabled: true
readinessState:
enabled: true
db:
enabled: true
redis:
enabled: true
custom:
enabled: true

方案二:配置告警规则

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
# prometheus/rules/service-alerts.yml
groups:
- name: service_health
rules:
- alert: ServiceDown
expr: up == 0
for: 2m
labels:
severity: critical
annotations:
summary: "服务 {{ $labels.instance }} 已宕机"

- alert: HighErrorRate
expr: rate(http_requests_total{status=~"5.."}[5m]) > 0.1
for: 5m
labels:
severity: warning
annotations:
summary: "服务 {{ $labels.instance }} 错误率过高"

3. 网络问题处理

对于跨境网络延迟问题,建议:

  1. 增加超时时间

    1
    2
    3
    4
    5
    # Prometheus 抓取配置
    scrape_configs:
    - job_name: '海外服务'
    scrape_timeout: 30s # 增加超时时间
    scrape_interval: 2m # 减少检查频率
  2. 添加网络质量监控

    1
    2
    3
    4
    5
    - alert: HighLatency
    expr: histogram_quantile(0.95, rate(http_request_duration_seconds_bucket[5m])) > 5
    for: 10m
    labels:
    severity: warning

一键排查脚本

当遇到类似问题时,可以使用以下脚本进行快速排查:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
#!/bin/bash

# 服务健康快速检查脚本

echo "========== 服务状态检查 =========="

# 检查 Prometheus targets
curl -s http://<prometheus-addr>:9090/api/v1/targets | \
jq '.data.activeTargets[] | select(.health=="down") | {job: .labels.job, instance: .labels.instance, error: .lastError}'

echo ""
echo "========== 错误率检查 =========="

# 检查最近5分钟错误率
curl -s "http://<prometheus-addr>:9090/api/v1/query?query=rate(http_server_requests_seconds_count{status=~'5..'}[5m])" | \
jq '.data.result[] | {instance: .metric.instance, error_rate: .value[1]}'

echo ""
echo "========== 响应时间检查 =========="

# 检查P99响应时间
curl -s "http://<prometheus-addr>:9090/api/v1/query?query=histogram_quantile(0.99, rate(http_server_requests_seconds_bucket[5m]))" | \
jq '.data.result[] | {instance: .metric.instance, p99: .value[1]}'

常见问题解答

Q:为什么监控显示 UP 但服务实际不可用?

A:可能的原因包括:健康检查端点只检查进程存活而不检查业务逻辑;负载均衡将请求路由到了少数正常实例;服务返回了 HTTP 200 但业务处理失败。建议增加更全面的健康检查和业务指标监控。

Q:如何快速定位是服务问题还是网络问题?

A:先在服务端本地测试服务是否正常(curl localhost:端口),如果本地正常但远程不行,那基本是网络问题。如果本地也不正常,那就是服务本身的问题。

Q:跨境服务响应慢应该如何优化?

A:可以考虑以下方案:1)增加监控超时时间;2)使用 CDN 加速静态资源;3)在当地部署服务副本;4)实施多区域容灾策略。

Q:Prometheus 告警阈值应该如何设置?

A:建议根据业务的实际负载和历史数据进行调优。初始可以设置较宽松的阈值,观察一段时间后再逐步收紧。对于关键服务,优先设置”服务不可用”告警,其次再考虑性能告警。

Q:如何避免监控”虚假绿色”的问题?

A:1)增加业务层面的健康检查,不仅检查进程,还要检查核心业务逻辑;2)配置服务依赖检查,确保下游服务不可用时能及时发现;3)定期进行故障演练,验证监控的有效性。

经验总结

  1. 监控需要分层次:基础设施监控、应用监控、业务监控,每一层都很重要
  2. 健康检查要全面:不能只看进程是否在跑,还要看业务是否正常
  3. 告警要有策略:区分关键告警和次要告警,避免告警疲劳
  4. 网络问题要耐心:跨境网络问题很多时候只能等待,别急着半夜起来改配置

参考资料


作者:小六,一个在上海努力搬砖的程序员

本文阅读量 --
Author:Margrop
Link:https://blog.margrop.com/post/2026-03-09-prometheus-service-up-but-not-responding/
版权声明:本文采用 CC BY-NC-SA 3.0 CN 协议进行许可