Skip to content

DX-596-together-gpu-clusters: sync with mintlify-docs#1277 - #62

Open
zainhas wants to merge 1 commit into
mainfrom
docs-sync/together-gpu-clusters/mintlify-docs-pr-1277
Open

zainhas wants to merge 1 commit into
mainfrom
docs-sync/together-gpu-clusters/mintlify-docs-pr-1277

Conversation

@zainhas

@zainhas zainhas commented Jul 20, 2026

Copy link
Copy Markdown
Collaborator

Syncs the together-gpu-clusters skill with togethercomputer/mintlify-docs#1277 ([TCL-7655] docs: update GPU cluster targeted scale-down guidance, merged commit 9c02409).

Changed docs files that triggered this sync

  • docs/gpu-clusters-management.mdx — restructured the Targeted scale-down section to document the new cluster-operator behavior deployed to prod.

Skill changes

  • Updated skills/together-gpu-clusters/references/cluster-management.md, "Targeted Scale-down" section:
    • Added the annotation-based approach (kubectl annotate node <node_name> node.together.ai/delete-node-on-scale-down=true) as the preferred Kubernetes marking method.
    • Kept kubectl cordon as an alternative for cases where you also want to block new pods immediately.
    • Kept the Slurm scontrol ... State=drain command for Slurm clusters.
    • Noted the Together Cloud UI Node is cordon/draining checkpoint that customers should wait for before triggering scale-down.
    • Added the recommendation to scale down one node at a time so the operator can drain each cleanly.
    • Called out that annotated and cordoned nodes are prioritized for deletion above all others.

No changes to SKILL.md, scripts, or other references — the update only affects operational Kubernetes/Slurm guidance in the cluster-management reference.

Generated by the Sync Skills Cursor Automation. Please review before merging.

Sync the together-gpu-clusters cluster-management reference with the
merged docs update. The Together cluster operator now supports an
annotation-driven targeted scale-down flow in addition to cordoning:

- Add `node.together.ai/delete-node-on-scale-down=true` annotation as
  the preferred Kubernetes marking method.
- Keep `kubectl cordon` as an alternative that also blocks new pods.
- Keep Slurm `scontrol ... State=drain` for Slurm clusters.
- Note the UI 'Node is cordon/draining' checkpoint before triggering
  scale-down, and the recommendation to scale one node at a time.
@zainhas zainhas self-assigned this Jul 20, 2026
@broly-code-security-scanner

Copy link
Copy Markdown

Broly Security Scan

Note

✅ Clean scan
No vulnerabilities detected in this PR.

Note

Re-scan this PR anytime with /broly scan — useful after /broly undismiss, or to refresh findings without a new push.

Broly — SAST (zai-org/GLM-5.2) · Secrets · SCA · GH Actions (zizmor) · Containers · SBOM · Powered by Together AI

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants