751-800 of 10000 results (37ms)
2026-07-28 ยง
10:34 <jiji@cumin1003> START - Cookbook sre.hosts.move-vlan for host wikikube-worker1137 [production]
10:34 <klausman@cumin1003> START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2001.codfw.wmnet [production]
10:34 <klausman@cumin1003> START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-codfw [production]
10:34 <jiji@cumin1003> START - Cookbook sre.hosts.reimage for host wikikube-worker1137.eqiad.wmnet with OS trixie [production]
10:28 <jforrester@deploy1003> jforrester: Continuing with deployment [production]
10:27 <jforrester@deploy1003> jforrester: Backport for [[gerrit:1307506|logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [production]
10:25 <jforrester@deploy1003> Started scap sync-world: Backport for [[gerrit:1307506|logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] [production]
10:21 <btullis@deploy1003> helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply [production]
10:21 <btullis@deploy1003> helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply [production]
10:21 <jgiannelos@deploy1003> helmfile [staging] DONE helmfile.d/services/kartotherian: apply [production]
10:20 <jiji@cumin1003> END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1137.eqiad.wmnet [production]
10:20 <jgiannelos@deploy1003> helmfile [staging] START helmfile.d/services/kartotherian: apply [production]
10:20 <jiji@cumin1003> START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1137.eqiad.wmnet [production]
10:20 <jiji@cumin1003> START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1137.eqiad.wmnet [production]
10:16 <cwilliams@cumin1003> END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2163: Maintenance [production]
09:37 <klausman@cumin1003> END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-staging-worker [production]
09:37 <klausman@cumin1003> END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet [production]
09:37 <klausman@cumin1003> START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet [production]
09:32 <kevinbazira@deploy1003> helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' . [production]
09:31 <klausman@cumin1003> END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet [production]
09:30 <ayounsi@cumin1003> END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool drmrs [reason: router upgrade, T431749] [production]
09:30 <cwilliams@cumin1003> START - Cookbook sre.mysql.pool pool db2163: Maintenance [production]
09:30 <ayounsi@cumin1003> START - Cookbook sre.dns.admin DNS admin: pool drmrs [reason: router upgrade, T431749] [production]
09:30 <kevinbazira@deploy1003> helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' . [production]
09:30 <ozge@deploy1003> helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' . [production]
09:22 <klausman@cumin1003> END (ERROR) - Cookbook sre.ganeti.reboot-vm (exit_code=97) for VM ml-serve-ctrl2001.codfw.wmnet [production]
09:22 <klausman@cumin1003> START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet [production]
09:22 <cmooney@cumin1003> END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d8-eqiad [production]
09:22 <cmooney@cumin1003> START - Cookbook sre.network.tls for network device lsw1-d8-eqiad [production]
09:21 <klausman@cumin1003> START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet [production]
09:20 <klausman@cumin1003> END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2002.codfw.wmnet [production]
09:20 <klausman@cumin1003> START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2002.codfw.wmnet [production]
09:19 <cmooney@cumin1003> END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-codfw [production]
09:18 <cmooney@cumin1003> START - Cookbook sre.network.tls for network device lsw1-f2-codfw [production]
09:18 <cmooney@cumin1003> END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e4-codfw [production]
09:18 <root@cumin1003> END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2163: Maintenance [production]
09:17 <cmooney@cumin1003> START - Cookbook sre.network.tls for network device lsw1-e4-codfw [production]
09:17 <cmooney@cumin1003> END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-codfw [production]
09:17 <cmooney@cumin1003> START - Cookbook sre.network.tls for network device lsw1-e2-codfw [production]
09:17 <cmooney@cumin1003> END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e5-codfw [production]
09:17 <cmooney@cumin1003> START - Cookbook sre.network.tls for network device lsw1-e5-codfw [production]
09:17 <cmooney@cumin1003> END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f4-codfw [production]
09:16 <cwilliams@cumin1003> END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2210: Maintenance [production]
09:16 <cmooney@cumin1003> START - Cookbook sre.network.tls for network device lsw1-f4-codfw [production]
09:15 <cwilliams@cumin1003> END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2182: Maintenance [production]
09:14 <klausman@cumin1003> END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2002.codfw.wmnet [production]
09:12 <root@cumin1003> START - Cookbook sre.mysql.pool pool db2163: Maintenance [production]
09:11 <XioNoX> rebooting cr2-drmrs - T431749 [production]
09:10 <ayounsi@cumin1003> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr2-drmrs,cr2-drmrs IPv6,cr2-drmrs.mgmt with reason: router upgrade [production]
09:06 <XioNoX> draining cr2-drmrs - T431749 [production]