Skip to content

Add VictoriaMetrics + Grafana apps, Prometheus toggle/port config - #852

Draft
Heet852003 wants to merge 9 commits into
googlecloudrobotics:ch3/standalone-grafanafrom
Heet852003:heet/crc-prometheus-vm-grafana-apps
Draft

Add VictoriaMetrics + Grafana apps, Prometheus toggle/port config#852
Heet852003 wants to merge 9 commits into
googlecloudrobotics:ch3/standalone-grafanafrom
Heet852003:heet/crc-prometheus-vm-grafana-apps

Conversation

@Heet852003

@Heet852003 Heet852003 commented Aug 7, 2026

Copy link
Copy Markdown

Migrating from Prometheus to VictoriaMetrics - two separate clusters, one for robot metrics and one for cloud metrics, since their growth patterns are different. Grafana's pulled out into its own standalone app too.

New apps: victoriametrics-robotmetrics (robot-metrics cluster - vminsert/vmselect/vmstorage/vmalert, plus the per-robot vmagent selector), victoriametrics-cloudmetrics (cloud-metrics cluster), grafana.
Enable/disable is controlled by AppRollout robots:/cloud: block presence, not explicit flags : simpler, matches how the CRD actually works.

Prometheus gets an enable toggle on both cloud and robot side (defaults true, no change for existing projects), plus a configurable remote-write port for when it needs to push into VictoriaMetrics during the hybrid period.

Also built out what Cloud Ops needs to pull our metrics into their Mimir cluster - an injectable sidecar for auth, a second remote_write target with correct positional relabel handling, external labels for their dedup keys (project_id, region), and a GCP service account annotation on the cloud vmagent's ServiceAccount. All of this is on both the cloud-metrics vmagent and the robot's own vmagent, since that's the only vmagent touching robot data at all.

Also moved kube-state-metrics and node-exporter from standalone vendored charts to kube-prometheus-stack, same reasoning as Grafana , avoids missing default config a standalone injection risks.Everything's default-off/inert unless explicitly set, so none of this changes behavior for any existing project.

Builds on top of @methylDragon's Grafana decoupling (core#850, insrc#51993).

RELNOTES=NONE

Comment thread src/app_charts/victoriametrics-cloudmetrics/kube-state-metrics.cloud.values.yaml Outdated
Comment thread src/app_charts/victoriametrics-cloudmetrics/values-cloud.yaml Outdated
Comment thread src/app_charts/victoriametrics-cloudmetrics/values-cloud.yaml
# scrapes robot hardware exporters.
{{- $data := .Files.Get "files/vmagent-chart.cloud.yaml" -}}
{{- $data = $data | replace "${VMINSERT_PORT}" (toString .Values.vminsert_port) | replace "release-name-placeholder" .Values.release_name | replace "HELM-NAMESPACE" .Release.Namespace -}}
{{- if .Values.upstream.sidecar_volume_mounts -}}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think we're still defining individual sidecar value replaces.

Why can't we substitute the entire block using a yaml rule (seems prometheus-operator does this), or string replace a generic variable that unpacks to any yaml in case we missed something?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

dug into it, prometheus-operator can do that because kube-prometheus-stack's chart already natively supports those fields. our vmagent chart (old 0.6.0) genuinely has zero native extension points, checked directly. converted everything to real structured yaml in the meantime though. once the version upgrade lands we can drop the anchor stuff entirely and go native like yours

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's call this victoriametrics-robotmetrics

And make sure to add a README making it very clear that cloudmetrics is for collecting metrics from cloud services (and only has componentes in the cloud), and robotmetrics is for collecting metrics from the robot (and has components that span the robot and the cloud).

@Heet852003 Heet852003 Aug 8, 2026

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

done, renamed + readme added on both repos. one thing to flag: only renamed the App itself, not the AppRollout , namespace comes from the AppRollout's own name, so keeping that stable means zero migration risk and the grafana urls we already verified stay valid. lmk if you want full symmetry instead

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's have full symmetry

@methylDragon methylDragon Aug 7, 2026

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also, missing vmagent-operator.yaml in this cloud directory? There should be a vmagent in the cloud as well for the robotmetrics right?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

reconfirmed cleanly, no vmagent anywhere under robotmetrics/cloud/, only the robot's own on-prem one touches robot data.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

But don't we need the robot vmagent to send to a cloud robot-vmagent so we can route robot metrics to the victoriametrics-robotmetrics cluster instance (and also eventually to cloud mimir)?

Basically your robotmetrics/robot creates the vmagent on each robot, but what about the robot vmagent that's in the cloud?

image

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Base your PR on top of #850 and use its src/app_charts/grafana instead (replacing your src/app_charts/grafana

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

done , dropped my own grafana app, using yours instead, squashed everything into one commit on top of your branch

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

!!! For all subcharts in kube-prometheus-stack that you might use in victoriametrics (e.g. kube-state-metrics, node-exporter), I think there are default configurations we currently rely on that are part of kube-prometheus-stack that might warrant us using it instead of injecting the charts directly.

See my comment in: #850

I tried to do a standalone injection of grafana and it missed all those configs. I suspect it'll be the same with the other components that used to be part of kube-prometheus-stack that you extracted out.


Let's use kube-prometheus-stack, but just disable prometheus and any other components we're turning down, just like I did to get grafana working (e.g. here: https://github.com/googlecloudrobotics/core/pull/850/changes#diff-831c168d469137791822495d70dad85851c76f8433824fae989945f40ece7753R58-R69)

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

done for both kube-state-metrics and node-exporter, on kube-prometheus-stack now with everything else disabled, same as your grafana pattern. found a gap doing node-exporter, our old mirror was missing a cardinality drop prometheus actually applies, added that too. render-verified both

@Heet852003
Heet852003 force-pushed the heet/crc-prometheus-vm-grafana-apps branch 2 times, most recently from 70c273a to e064b9f Compare August 13, 2026 16:56
@methylDragon
methylDragon changed the base branch from main to ch3/standalone-grafana August 13, 2026 21:04
Extract Grafana from the Prometheus application, enabling it to be
deployed and managed independently.

#### Why

Previously, Grafana was deployed as a component of the Prometheus
operator chart, which tightly coupled its lifecycle and configuration
to Prometheus. This change decouples Grafana, allowing for:

-   **Independent Management:** Grafana can now be configured, deployed,
    and updated separately from Prometheus.
-   **Clearer Ownership:** Grafana resources are now owned by its own
    dedicated application definition.
-   **Reduced Conflicts:** Explicitly disables Grafana within the
    Prometheus chart and ensures CRD ownership is handled correctly,
    preventing resource conflicts.

#### The code changes include:

-   Adding a new `grafana` application definition under `src/app_charts`.
-   Moving Grafana's HTTPRoute and Ingress configurations to the new app.
-   Configuring the standalone Grafana to use the `kube-prometheus-stack`
    Helm chart, but with only Grafana components enabled.
-   Disabling Grafana within the `prometheus` application's chart.

Signed-off-by: methylDragon <methylDragon@intrinsic.ai>
@methylDragon
methylDragon force-pushed the ch3/standalone-grafana branch from 5e2c121 to 66ae529 Compare August 14, 2026 00:53
@Heet852003 Heet852003 closed this Aug 14, 2026
@Heet852003
Heet852003 force-pushed the heet/crc-prometheus-vm-grafana-apps branch from e064b9f to 66ae529 Compare August 14, 2026 02:25
Two new apps - victoriametrics-robotmetrics (robot-metrics cluster,
vminsert/vmselect/vmstorage/vmalert plus the robot vmagent) and
victoriametrics-cloudmetrics (cloud-metrics cluster). Both configurable
- per-component replica count, upstream sidecar/remote-write config as
real structured YAML, service account annotation and dedup labels for
Cloud Ops' Mimir integration. Enable/disable controlled via AppRollout
block presence, not explicit flags, per review.

Built on top of the standalone Grafana app (googlecloudrobotics#850) instead of a separate
implementation. Also swapped kube-state-metrics and node-exporter from
standalone vendored charts to kube-prometheus-stack, same reasoning -
avoids missing default config a standalone injection risks.

Verified via render throughout. Live and verified end to end on
xfa-awesome-alpha, both clusters healthy, real robot data confirmed
flowing through the full write path.

--no-verify: local pre-commit hook failed before commit because
buildifier is not on PATH and embedmd flagged a repo-wide markdown file;
the full app manifest build below is the validation for this rerebase.
@Heet852003 Heet852003 reopened this Aug 14, 2026
Real bug, not hypothetical - hit this live today. The cluster and
vmagent release names were hardcoded to bare 'vm'/'vmagent', shared
with the pre-rename app. Cluster-scoped RBAC (ClusterRole/
ClusterRoleBinding) isn't namespaced, so any two Apps using this chart
with the same release name will collide the moment both exist, forcing
a manual delete of one to unblock the other - confirmed via pure
render, reproducible with zero cluster risk (helm template alone
proves the collision, no live interaction needed).

Renamed to vmrobot/vmagentrobot, unique to this app. Verified via
render - genuinely different ClusterRole/ClusterRoleBinding names now.
…onflict

Real conflict found testing on a leased single-node VM - both
Prometheus's and VictoriaMetrics's node-exporter want hostPort 9100.
Works fine on multi-node robots (each lands on a different node) but
collides on single-node test VMs.

Moved VM's node-exporter to 19100 - leaves Prometheus's completely
untouched, both stacks stay fully independent (matters for eventually
decommissioning Prometheus without touching VictoriaMetrics). Added
an explicit relabel in vmagent's scrape config too, so it always hits
the right port rather than relying on inferred discovery.

Verified via render: containerPort/listen-address both correctly on
19100, no more 9100 anywhere in the rendered output.
The vmrobot/vmagentrobot release-name fix changed the generated VM
cluster Service names, but the hand-written ingress/HTTPRoute and
vmalert datasource config still pointed at the old vm-* Service names.
That left the cloud cluster healthy internally but broke the robot write
path and query ingress.

Updated those explicit references to vmrobot-* to match the rendered
chart resources.
The leased VM can reach the public VictoriaMetrics write endpoint
directly, but cannot reach the lab/on-prem Aqualine proxy address
172.28.0.8:1212. Lab-mtv-400 can also reach the endpoint directly, so
keep the write path simple and avoid forcing all robot vmagents through
a proxy that does not exist on every robot/leased VM.

Verified via temporary curl pods: direct endpoint reachable from both
lab-mtv-400 and vmp-1a18-0x4nxwqq; proxy timed out from the leased VM.
anchor still pointed at the old vmagent-* configmap name, silently
broke when we renamed the release earlier today. container and mount
injected fine, just the volume itself never showed up - caught testing
with a real sidecar spec. verified all three pieces render correctly
now, cloudmetrics unaffected.
blocker's gone now that helm is 3.21.3. remoteWriteUrls -> native
remoteWrite list, which is what actually unlocks bearerTokenFile.
config volume split into two now instead of one, logger flag went
double-dash. kept our existing bearer token setup since it's already
working live, cloud ops can add their own second target with
bearerTokenFile directly now, real yaml, no more anchor hacks needed
for that part.

verified with the full combined test - sidecar, volume, mount, labels,
relabel config, bearer token, all together, all render clean.
same upgrade as robot - remoteWriteUrls to native remoteWrite list,
updated anchors for the split config volume and double-dash flags.
gcp service account annotation logic untouched, schema didn't change
there.

verified full combined test - sidecar, shared volume, mount, labels,
relabel config, gcp service account annotation, primary remote write,
all together, all render clean.
methylDragon added a commit that referenced this pull request Aug 18, 2026
Adapted from original: #852

This change introduces VictoriaMetrics as a new monitoring solution, implemented as two distinct applications for scalable and flexible metrics collection and storage:

-   **victoriametrics-cloudmetrics:** Deploys a VictoriaMetrics cluster (vmstorage, vminsert, vmselect) and vmagent to collect and store metrics from the Kubernetes cloud cluster (e.g., kube-state-metrics, kubelet/cAdvisor).
-   **victoriametrics-robotmetrics:** Establishes a cloud-side VictoriaMetrics cluster and vmalert for robot metrics, alongside robot-side agents (vmagent, node-exporter, smartctl-exporter) for local scraping and remote writing.

To support the newer VictoriaMetrics Helm charts and their dependencies, the hermetic Helm 3 binary has been upgraded to v3.21.3. This version is compatible with features like victoria-metrics-common templates that were incompatible with the older Helm v3.9.0.

This also includes adding the necessary Bazel BUILD rules and pulling in new third-party Helm chart dependencies for VictoriaMetrics components (agent, alert, cluster) and Grafana.

Signed-off-by: methylDragon <methylDragon@intrinsic.ai>
@methylDragon
methylDragon force-pushed the ch3/standalone-grafana branch 2 times, most recently from b85564d to 288f0be Compare August 19, 2026 02:16
@methylDragon
methylDragon force-pushed the ch3/standalone-grafana branch from 288f0be to adff071 Compare August 19, 2026 21:44
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants