251-300 of 10000 results (37ms)
2026-07-23 ยง
13:34 <brouberol@deploy1003> helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply [production]
13:34 <brouberol@deploy1003> helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply [production]
13:33 <brouberol@deploy1003> helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply [production]
13:33 <brouberol@deploy1003> helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply [production]
13:33 <kharlan@deploy1003> kharlan, emc-wmf: Continuing with deployment [production]
13:31 <kharlan@deploy1003> kharlan, emc-wmf: Backport for [[gerrit:1311851|EventStreamConfig: remove stream used in past experiments (T428265)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [production]
13:30 <cmooney@cumin1003> START - Cookbook sre.dns.netbox [production]
13:28 <kharlan@deploy1003> Started scap sync-world: Backport for [[gerrit:1311851|EventStreamConfig: remove stream used in past experiments (T428265)]] [production]
13:17 <hashar@deploy1003> Finished deploy [integration/docroot@2199146]: build: License GPL2.0+ / updating npm dependencies (duration: 00m 14s) [production]
13:17 <hashar@deploy1003> Started deploy [integration/docroot@2199146]: build: License GPL2.0+ / updating npm dependencies [production]
13:14 <sukhe> sukhe@lvs1019:~$ sudo systemctl restart pybal.service [production]
13:07 <sukhe> sukhe@lvs1020:~$ sudo systemctl restart pybal.service [production]
12:58 <marostegui@cumin1003> END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: Repooling [production]
12:51 <godog> delete tool -- now unused [tools.toolschecker]
12:49 <root@cumin1003> START - Cookbook sre.mysql.pool pool db2196: Maintenance [production]
12:39 <cwilliams@cumin1003> dbctl commit (dc=all): 'Depooling db2196 (T431660)', diff saved to https://phabricator.wikimedia.org/P95087 and previous config saved to /var/cache/conftool/dbconfig/20260723-123952-cwilliams.json [production]
12:39 <cwilliams@cumin1003> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2196.codfw.wmnet with reason: Maintenance [production]
12:36 <root@cumin1003> END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2191: Maintenance [production]
12:30 <dcaro@cloudcumin1001> END (PASS) - Cookbook wmcs.toolforge.component.deploy (exit_code=0) for component jobs-cli [tools]
12:29 <brouberol@deploy1003> helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. [production]
12:28 <brouberol@deploy1003> helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. [production]
12:22 <dcaro@cloudcumin1001> START - Cookbook wmcs.toolforge.component.deploy for component jobs-cli [tools]
12:20 <dcaro@cloudcumin1001> END (PASS) - Cookbook wmcs.toolforge.component.deploy (exit_code=0) for component jobs-cli [toolsbeta]
12:13 <ozge@deploy1003> helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . [production]
12:13 <marostegui@cumin1003> START - Cookbook sre.mysql.pool pool db2207: Repooling [production]
12:12 <marostegui@cumin1003> END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2207: Repooling [production]
12:12 <marostegui@cumin1003> START - Cookbook sre.mysql.pool pool db2207: Repooling [production]
12:12 <dcaro@cloudcumin1001> START - Cookbook wmcs.toolforge.component.deploy for component jobs-cli [toolsbeta]
12:10 <taavi@cloudcumin1001> END (PASS) - Cookbook wmcs.vps.remove_instance (exit_code=0) for instance metricsinfra-prometheus-3 [metricsinfra]
12:09 <taavi@cloudcumin1001> START - Cookbook wmcs.vps.remove_instance for instance metricsinfra-prometheus-3 [metricsinfra]
12:05 <taavi@cloudcumin1001> END (PASS) - Cookbook wmcs.vps.refresh_puppet_certs (exit_code=0) on metricsinfra-prometheus-5.metricsinfra.eqiad1.wikimedia.cloud [metricsinfra]
12:04 <taavi@cloudcumin1001> START - Cookbook wmcs.vps.refresh_puppet_certs on metricsinfra-prometheus-5.metricsinfra.eqiad1.wikimedia.cloud [metricsinfra]
12:03 <taavi@cloudcumin1001> END (FAIL) - Cookbook wmcs.vps.refresh_puppet_certs (exit_code=99) on metricsinfra-prometheus-5.metricsinfra.eqiad1.wikimedia.cloud [metricsinfra]
12:02 <taavi@cloudcumin1001> START - Cookbook wmcs.vps.refresh_puppet_certs on metricsinfra-prometheus-5.metricsinfra.eqiad1.wikimedia.cloud [metricsinfra]
12:00 <dcaro@cloudcumin1001> END (PASS) - Cookbook wmcs.toolforge.component.deploy (exit_code=0) for component jobs-api [tools]
11:56 <marostegui@cumin1003> END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2235.codfw.wmnet with OS trixie [production]
11:50 <root@cumin1003> START - Cookbook sre.mysql.pool pool db2191: Maintenance [production]
11:47 <dcaro@cloudcumin1001> START - Cookbook wmcs.toolforge.component.deploy for component jobs-api [tools]
11:46 <jiji@cumin1003> END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1070.eqiad.wmnet [production]
11:46 <jiji@cumin1003> END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1070.eqiad.wmnet [production]
11:46 <jiji@cumin1003> START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1070.eqiad.wmnet [production]
11:46 <dcaro@cloudcumin1001> END (PASS) - Cookbook wmcs.toolforge.component.deploy (exit_code=0) for component jobs-api [toolsbeta]
11:43 <cwilliams@cumin1003> dbctl commit (dc=all): 'Depooling db2191 (T431660)', diff saved to https://phabricator.wikimedia.org/P95080 and previous config saved to /var/cache/conftool/dbconfig/20260723-114308-cwilliams.json [production]
11:43 <cwilliams@cumin1003> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2191.codfw.wmnet with reason: Maintenance [production]
11:41 <root@cumin1003> END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2186: Maintenance [production]
11:35 <dcaro@cloudcumin1001> START - Cookbook wmcs.toolforge.component.deploy for component jobs-api [toolsbeta]
11:35 <cmooney@cumin1003> END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 46375 [production]
11:34 <cmooney@cumin1003> START - Cookbook sre.network.peering with action 'configure' for AS: 46375 [production]
11:33 <marostegui@cumin1003> END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2235.codfw.wmnet with reason: host reimage [production]
11:28 <marostegui@cumin1003> START - Cookbook sre.hosts.downtime for 2:00:00 on db2235.codfw.wmnet with reason: host reimage [production]