> ## Documentation Index
> Fetch the complete documentation index at: https://docs.coreweave.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Visualize CoreWeave infrastructure alerts

> View CoreWeave infrastructure alerts such as GPU failures and thermal violations on your W&B experiment run plots.

W\&B에 로깅하는 머신러닝 실험 중에 발생한 GPU 장애, 열 임계값 위반 등의 인프라 알림을 확인할 수 있습니다. 지원되는 [CoreWeave Kubernetes Service (CKS)](/products/cks) cluster에서 실행하면서 이 인테그레이션을 활성화하고 이 페이지의 사전 요구 사항을 충족하면, [CoreWeave Mission Control](https://www.coreweave.com/mission-control)이 [W\&B run](/ko/products/wandb/runs) 동안 컴퓨팅 인프라를 모니터링합니다.

<Note>
  이 기능은 Preview 단계입니다. 액세스 권한이 필요하면 W\&B 담당자에게 문의하세요.
</Note>

<h2 id="prerequisites">
  사전 요구 사항
</h2>

이 인테그레이션이 전체 과정에서 제대로 작동하려면 다음 조건을 충족해야 합니다.

| 사전 요구 사항 | 세부 정보 |
| - | - |
| **CoreWeave 플랫폼** | [CoreWeave Kubernetes Service (CKS)](/products/cks) cluster에서만 사용 가능합니다. CoreWeave 베어 메탈 cluster나 CoreWeave Classic에서는 사용 가능하지 않습니다. CKS에서 [SUNK](/products/sunk)를 통해 실행하는 트레이닝 작업도 이 요구 사항을 충족합니다. |
| **W\&B Python SDK** | 트레이닝 작업에서 run을 로깅할 때는 `wandb` 패키지 버전 `0.20.1` 이상을 사용하세요. |
| **W\&B Server (Dedicated Cloud 또는 Self-Managed)** | W\&B Dedicated Cloud 또는 W\&B Self-Managed 배포를 사용하는 경우 W\&B Server 버전 `0.73.0` 이상을 사용하세요. 서버가 CoreWeave 관측성 데이터를 수신할 수 있도록 W\&B 앱 파드에 `SERVER_FLAG_ENABLE_CORE_WEAVE_OBSERVABILITY` 환경 변수를 설정하세요. |

오류가 발생하면 CoreWeave가 해당 정보를 W\&B로 전송하고, W\&B는 프로젝트 workspace의 run 플롯에 인프라 정보를 표시합니다. CoreWeave는 일부 문제를 자동으로 해결하려고 시도하며, W\&B는 그 내용을 run 페이지에 표시합니다.


## Related topics

- [Introduction to CoreWeave Alerts](/platform/coreweave-alerts.md)
