351-400 of 10000 results (130ms)
2026-07-28 ยง
11:14 <klausman@cumin1003> END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2003.codfw.wmnet [production]
11:11 <cwilliams@cumin1003> START - Cookbook sre.mysql.pool pool db2219: Maintenance [production]
11:10 <root@cumin1003> END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2219: Maintenance [production]
11:04 <root@cumin1003> START - Cookbook sre.mysql.pool pool db2219: Maintenance [production]
11:04 <klausman@cumin1003> START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2003.codfw.wmnet [production]
11:04 <root@cumin1003> START - Cookbook sre.mysql.pool pool db2164: Maintenance [production]
11:03 <klausman@cumin1003> END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet [production]
11:03 <klausman@cumin1003> START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet [production]
11:03 <root@cumin1003> START - Cookbook sre.mysql.pool pool db2208: Maintenance [production]
10:57 <cwilliams@cumin1003> dbctl commit (dc=all): 'Depooling db2164 (T431660)', diff saved to https://phabricator.wikimedia.org/P95332 and previous config saved to /var/cache/conftool/dbconfig/20260728-105749-cwilliams.json [production]
10:57 <cwilliams@cumin1003> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2164.codfw.wmnet with reason: Maintenance [production]
10:57 <cwilliams@cumin1003> dbctl commit (dc=all): 'Depooling db2208 (T431660)', diff saved to https://phabricator.wikimedia.org/P95331 and previous config saved to /var/cache/conftool/dbconfig/20260728-105711-cwilliams.json [production]
10:57 <cwilliams@cumin1003> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2208.codfw.wmnet with reason: Maintenance [production]
10:56 <cwilliams@cumin1003> dbctl commit (dc=all): 'Depooling db2219 (T431660)', diff saved to https://phabricator.wikimedia.org/P95330 and previous config saved to /var/cache/conftool/dbconfig/20260728-105652-cwilliams.json [production]
10:56 <cwilliams@cumin1003> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2219.codfw.wmnet with reason: Maintenance [production]
10:53 <jiji@cumin1003> START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1137 - jiji@cumin1003" [production]
10:52 <klausman@cumin1003> END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet [production]
10:47 <klausman@cumin1003> START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet [production]
10:47 <klausman@cumin1003> END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2001.codfw.wmnet [production]
10:47 <klausman@cumin1003> START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2001.codfw.wmnet [production]
10:39 <klausman@cumin1003> END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2001.codfw.wmnet [production]
10:35 <jiji@cumin1003> START - Cookbook sre.dns.netbox [production]
10:34 <jforrester@deploy1003> Finished scap sync-world: Backport for [[gerrit:1307506|logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] (duration: 09m 31s) [production]
10:34 <jiji@cumin1003> START - Cookbook sre.hosts.move-vlan for host wikikube-worker1137 [production]
10:34 <klausman@cumin1003> START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2001.codfw.wmnet [production]
10:34 <klausman@cumin1003> START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-codfw [production]
10:34 <jiji@cumin1003> START - Cookbook sre.hosts.reimage for host wikikube-worker1137.eqiad.wmnet with OS trixie [production]
10:28 <jforrester@deploy1003> jforrester: Continuing with deployment [production]
10:27 <jforrester@deploy1003> jforrester: Backport for [[gerrit:1307506|logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [production]
10:25 <jforrester@deploy1003> Started scap sync-world: Backport for [[gerrit:1307506|logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] [production]
10:21 <btullis@deploy1003> helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply [production]
10:21 <btullis@deploy1003> helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply [production]
10:21 <jgiannelos@deploy1003> helmfile [staging] DONE helmfile.d/services/kartotherian: apply [production]
10:20 <jiji@cumin1003> END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1137.eqiad.wmnet [production]
10:20 <jgiannelos@deploy1003> helmfile [staging] START helmfile.d/services/kartotherian: apply [production]
10:20 <jiji@cumin1003> START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1137.eqiad.wmnet [production]
10:20 <jiji@cumin1003> START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1137.eqiad.wmnet [production]
10:16 <cwilliams@cumin1003> END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2163: Maintenance [production]
09:37 <klausman@cumin1003> END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-staging-worker [production]
09:37 <klausman@cumin1003> END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet [production]
09:37 <klausman@cumin1003> START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet [production]
09:32 <kevinbazira@deploy1003> helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . [production]
09:31 <klausman@cumin1003> END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet [production]
09:30 <ayounsi@cumin1003> END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool drmrs [reason: router upgrade, T431749] [production]
09:30 <cwilliams@cumin1003> START - Cookbook sre.mysql.pool pool db2163: Maintenance [production]
09:30 <ayounsi@cumin1003> START - Cookbook sre.dns.admin DNS admin: pool drmrs [reason: router upgrade, T431749] [production]
09:30 <kevinbazira@deploy1003> helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' . [production]
09:30 <ozge@deploy1003> helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . [production]
09:22 <klausman@cumin1003> END (ERROR) - Cookbook sre.ganeti.reboot-vm (exit_code=97) for VM ml-serve-ctrl2001.codfw.wmnet [production]
09:22 <klausman@cumin1003> START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet [production]