Google Vertex AI Endpoint metrics
A Vertex AI endpoint serves online predictions from one or more deployed models. Ensure your cloud platform is configured in SolarWinds Observability SaaS to collect this service's data. See Add a GCP cloud account.
Many of the collected metrics from Vertex AI Endpoint entities are displayed as widgets in SolarWinds Observability explorers; additional metrics may be collected and available in the Metrics Explorer. You can also create an alert for when an entity's metric value moves out of a specific range. See Entities in SolarWinds Observability SaaS for information about entity types in SolarWinds Observability SaaS.
The following table lists aiplatform.googleapis.com/prediction/online in the search box.
| Metric | Unit | Description |
|---|---|---|
sw.metrics.healthscore
|
Percent (%) |
Health state. The health state provides real-time insight into the overall health and performance of your monitored entities. The health state is determined based on anomalies detected for the entity, alerts triggered for the entity's metrics, and the status of the entity. The health state is displayed as one of the following four states and colors: Good, Moderate, Bad, or Unknown. You can determine the impact of the alerts, anomalies, and statuses on the health of an entity type by going to Settings > Health, and selecting a specific entity type. You can also customize the impact. To view the health of Google Vertex AI Endpoint entities in the Metrics Explorer, filter the |
gcp.aiplatform.googleapis.com.
|
Count per second | Number of online predictions served. |
gcp.aiplatform.googleapis.com.
|
Milliseconds (ms) | Online prediction latency distribution. |
gcp.aiplatform.googleapis.com.
|
Count per second | Number of online prediction errors. |
gcp.aiplatform.googleapis.com.
|
Count per second | Number of online prediction responses, by response code. |
gcp.aiplatform.googleapis.com.
|
Scaled percentage | Fraction of the CPU allocated for online prediction that is currently in use. |
gcp.aiplatform.googleapis.com.
|
Bytes | Amount of the memory allocated for online prediction that is currently in use. |
gcp.aiplatform.googleapis.com.
|
Scaled percentage | Average fraction of time that the accelerators were actively processing for the endpoint. |
gcp.aiplatform.googleapis.com.
|
Bytes | Amount of accelerator memory allocated and in use for online prediction. |
gcp.aiplatform.googleapis.com.
|
Bytes per second | Number of bytes received over the network by the deployed model replica. |
gcp.aiplatform.googleapis.com.
|
Bytes per second | Number of bytes sent over the network by the deployed model replica. |
gcp.aiplatform.googleapis.com.
|
Milliseconds (ms) | Online prediction latency of the private deployed model. |
gcp.aiplatform.googleapis.com.
|
Count per second | Number of online prediction responses from the private deployed model. |