Development Tool: Environment Inspection and Diagnostics
Health checks, log diagnosis, alarm interpretation, and troubleshooting for CloudBase environments — covering typical problems such as 429 rate limiting, cloud function 404, and zero invocations
How to Use
See How to Use Skill for detailed usage.
Test Skill
You can use the following prompts to test:
- "Run a health inspection on my CloudBase environment"
- "Cloud function invocations suddenly dropped to zero — help me find out why"
- "Interpret this CPU utilization alarm and suggest what to do about it"
Use AI to inspect environments, diagnose logs and alarms
Installation and Viewing
To install all CloudBase Skills, run:
npx skills add tencentcloudbase/cloudbase-skills
To install only the current Skill, run:
npx skills add https://github.com/tencentcloudbase/skills --skill ops-inspector
View current Skill online: ops-inspector
Skill Rules Original Text
View SKILL.md Original
## Sibling skills (local only)
Sibling CloudBase skills ship beside this skill. Use local relative paths such as `../auth-tool-cloudbase/SKILL.md`.
If a referenced sibling skill file is missing from this environment, ask the user to install the full CloudBase plugin (or the missing skill). Do **not** HTTP-fetch remote skill or protocol markdown into the agent context.
## Activation Contract
### Use this first when
- The user wants to check the health or status of CloudBase resources (cloud functions, CloudRun, databases, storage, etc.).
- The user reports errors, failures, or abnormal behavior and wants a quick diagnosis.
- The user asks for an "inspection", "health check", "巡检", "诊断", or "troubleshooting" of their CloudBase environment.
- The user wants to review recent error logs across services.
- The user asks **告警解读** questions: whether a **CPU 告警** is normal, what **峰值 QPS** was, or whether throttle/error metrics look healthy.
- The symptom matches a v3 fault playbook: **429 / 限频**, **云函数 404**, **ACCESS_TOKEN_INVALID**, or **调用量为 0**.
### Read before writing code if
- The inspection reveals code-level issues in cloud functions or CloudRun services — then read the relevant implementation skill before suggesting fixes.
- The user wants to fix a problem found during inspection rather than just diagnose it.
### Then also read
- Alarm interpretation baselines -> `references/alarm-interpretation.md`
- Fault playbooks (429 / 404 / token / zero calls) -> `references/fault-playbooks.md`
- Cloud function issues -> `../cloud-functions/SKILL.md`
- CloudRun issues -> `../cloudrun-development/SKILL.md`
- Database issues -> `../postgresql-development-cloudbase/SKILL.md` for CloudBase PG / PostgreSQL, `../relational-database-mcp-cloudbase/SKILL.md` for MySQL, or `../cloudbase-document-database-web-sdk/SKILL.md` for NoSQL
- Auth readiness (token failures) -> `../auth-tool-cloudbase/SKILL.md`
- Platform overview -> `../cloudbase-platform/SKILL.md`
### Do NOT use for
- Deploying new resources or writing application code. This skill is read-only and diagnostic.
- Replacing proper monitoring/alerting infrastructure. It provides point-in-time inspection, not continuous monitoring.
- Directly fixing problems — it diagnoses and recommends; actual fixes should use the appropriate implementation skill.
- Fetching metrics by guessing cloud API Actions. **Never** use `callCloudApi` for monitor curves — always use `queryEnv(action="metrics")`.
### Common mistakes / gotchas
- Running a full inspection without first confirming the environment is bound (`auth` tool must show logged-in and env-bound state).
- Ignoring CLS log service status — if CLS is not enabled, `queryLogs` will fail; always check first with `queryLogs(action="checkLogService")`.
- Searching logs without a time range — this can return excessive or irrelevant results. Always scope searches to a relevant time window.
- Treating a single error log as the root cause without correlating across resources. A function error may stem from a database or config issue.
- Answering "峰值 QPS" / "CPU 告警是否正常" from screenshots or memory instead of `queryEnv(action="metrics")`.
- Calling `callCloudApi` with invented `GetMonitorData` / `DescribeCurveData` parameters — the metrics branch already wraps Manager SDK.
### Minimal checklist
- [ ] Environment is bound and accessible (`envQuery(action="info")`)
- [ ] Metrics pulled with `queryEnv(action="metrics")` when the question involves QPS / CPU / throttle / invocation volume
- [ ] CLS log service is enabled (`queryLogs(action="checkLogService")`) when log diagnosis is needed
- [ ] Matching fault playbook selected when symptoms match 429 / function 404 / ACCESS_TOKEN_INVALID / 调用量为 0
- [ ] Time range is specified for any log or metrics searches
- [ ] Findings are summarized with severity levels, **告警解读**, and actionable recommendations
---
## How to use this skill (for a coding agent)
### Ops Inspector v3 additions
v3 adds two mandatory capabilities on top of log/resource inspection:
1. **告警解读** — pull metrics, compare to baselines in `references/alarm-interpretation.md`, answer CPU-alert / peak-QPS style questions in plain language.
2. **