Google Vertex AI Index Endpoint metrics
A Vertex AI index endpoint serves nearest-neighbor queries against one or more deployed vector indexes. Ensure your cloud platform is configured in SolarWinds Observability SaaS to collect this service's data. See Add a GCP cloud account.
Many of the collected metrics from Vertex AI Index Endpoint entities are displayed as widgets in SolarWinds Observability explorers; additional metrics may be collected and available in the Metrics Explorer. You can also create an alert for when an entity's metric value moves out of a specific range. See Entities in SolarWinds Observability SaaS for information about entity types in SolarWinds Observability SaaS.
The following table lists aiplatform.googleapis.com/matching_engine in the search box.
| Metric | Unit | Description |
|---|---|---|
sw.metrics.healthscore
|
Percent (%) |
Health state. The health state provides real-time insight into the overall health and performance of your monitored entities. The health state is determined based on anomalies detected for the entity, alerts triggered for the entity's metrics, and the status of the entity. The health state is displayed as one of the following four states and colors: Good, Moderate, Bad, or Unknown. You can determine the impact of the alerts, anomalies, and statuses on the health of an entity type by going to Settings > Health, and selecting a specific entity type. You can also customize the impact. To view the health of Google Vertex AI Index Endpoint entities in the Metrics Explorer, filter the |
gcp.aiplatform.googleapis.com.matching_engine.query.requestCount
|
Count per second | Number of nearest-neighbor query requests to the index endpoint. |
gcp.aiplatform.googleapis.com.matching_engine.query.latencies
|
Milliseconds (ms) | Latency distribution for nearest-neighbor query requests to the index endpoint. |
gcp.aiplatform.googleapis.com.matching_engine.cpu.requestUtilization
|
Scaled percentage | Fraction of the CPU request capacity currently used by the index endpoint. |
gcp.aiplatform.googleapis.com.matching_engine.memory.usedBytes
|
Bytes | Memory used by the deployed index on the index endpoint. |
gcp.aiplatform.googleapis.com.matching_engine.currentReplicas
|
Count | Current number of active replicas serving the deployed index. |
gcp.aiplatform.googleapis.com.matching_engine.currentShards
|
Count | Current number of shards in the deployed index. |