Documentation forSolarWinds Observability SaaS

Google Vertex AI Endpoint metrics

A Vertex AI endpoint serves online predictions from one or more deployed models. Ensure your cloud platform is configured in SolarWinds Observability SaaS to collect this service's data. See Add a GCP cloud account.

Many of the collected metrics from Vertex AI Endpoint entities are displayed as widgets in SolarWinds Observability explorers; additional metrics may be collected and available in the Metrics Explorer. You can also create an alert for when an entity's metric value moves out of a specific range. See Entities in SolarWinds Observability SaaS for information about entity types in SolarWinds Observability SaaS.

The following table lists some of the metrics collected for these entities. To see the Vertex AI Endpoint metrics in the Metrics Explorer, type aiplatform.googleapis.com/prediction/online in the search box.

Metric Unit Description
sw.metrics.healthscore Percent (%)

Health state. The health state provides real-time insight into the overall health and performance of your monitored entities. The health state is determined based on anomalies detected for the entity, alerts triggered for the entity's metrics, and the status of the entity. The health state is displayed as one of the following four states and colors: Good, Moderate, Bad, or Unknown. You can determine the impact of the alerts, anomalies, and statuses on the health of an entity type by going to Settings > Health, and selecting a specific entity type. You can also customize the impact.

To view the health of Google Vertex AI Endpoint entities in the Metrics Explorer, filter the sw.metrics.healthscore metric by entity_types and select gcpvertexaiendpoint.

gcp.aiplatform.googleapis.com.
prediction.online.predictionCount
Count per second Number of online predictions served.
gcp.aiplatform.googleapis.com.
prediction.online.predictionLatencies
Milliseconds (ms) Online prediction latency distribution.
gcp.aiplatform.googleapis.com.
prediction.online.errorCount
Count per second Number of online prediction errors.
gcp.aiplatform.googleapis.com.
prediction.online.responseCount
Count per second Number of online prediction responses, by response code.
gcp.aiplatform.googleapis.com.
prediction.online.cpu.utilization
Scaled percentage Fraction of the CPU allocated for online prediction that is currently in use.
gcp.aiplatform.googleapis.com.
prediction.online.memory.bytesUsed
Bytes Amount of the memory allocated for online prediction that is currently in use.
gcp.aiplatform.googleapis.com.
prediction.online.accelerator.dutyCycle
Scaled percentage Average fraction of time that the accelerators were actively processing for the endpoint.
gcp.aiplatform.googleapis.com.
prediction.online.accelerator.memory.bytesUsed
Bytes Amount of accelerator memory allocated and in use for online prediction.
gcp.aiplatform.googleapis.com.
prediction.online.network.receivedBytesCount
Bytes per second Number of bytes received over the network by the deployed model replica.
gcp.aiplatform.googleapis.com.
prediction.online.network.sentBytesCount
Bytes per second Number of bytes sent over the network by the deployed model replica.
gcp.aiplatform.googleapis.com.
prediction.online.private.predictionLatencies
Milliseconds (ms) Online prediction latency of the private deployed model.
gcp.aiplatform.googleapis.com.
prediction.online.private.responseCount
Count per second Number of online prediction responses from the private deployed model.