1-50 of 10000 results (113ms)
2026-08-06 ยง
16:50 <brouberol@cumin1003> END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1002.eqiad.wmnet with OS bookworm [production]
16:35 <brouberol@cumin1003> END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1001.eqiad.wmnet with OS bookworm [production]
16:19 <brouberol@cumin1003> END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage [production]
16:16 <brouberol@cumin1003> START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage [production]
16:05 <brouberol@cumin1003> END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage [production]
16:00 <brouberol@cumin1003> START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage [production]
15:57 <brouberol@cumin1003> START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm [production]
15:43 <brouberol@cumin1003> START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm [production]
15:29 <brouberol@cumin1003> END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm [production]
15:29 <brouberol@cumin1003> END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm [production]
14:58 <brouberol@cumin1003> START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm [production]
14:57 <cdobbins@cumin1003> END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs (T428495) [production]
14:55 <cdobbins@cumin1003> START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs (T428495) [production]
14:55 <brouberol@cumin1003> END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-ui1001.eqiad.wmnet with OS bookworm [production]
14:54 <brouberol@cumin1003> END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-presto1001.eqiad.wmnet with OS bookworm [production]
14:52 <cdobbins@cumin1003> END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru (T428495) [production]
14:49 <cdobbins@cumin1003> START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru (T428495) [production]
14:48 <brouberol@cumin1003> END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm [production]
14:46 <cdobbins@cumin1003> END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams (T428495) [production]
14:44 <cdobbins@cumin1003> START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams (T428495) [production]
14:43 <brouberol@cumin1003> START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm [production]
14:42 <brouberol@cumin1003> END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-worker1002.eqiad.wmnet with OS bookworm [production]
14:42 <brouberol@cumin1003> START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm [production]
14:42 <cdobbins@cumin1003> END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo (T428495) [production]
14:40 <cdobbins@cumin1003> START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo (T428495) [production]
14:40 <brouberol@cumin1003> END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm [production]
14:38 <brouberol@cumin1003> END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage [production]
14:34 <brouberol@cumin1003> END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage [production]
14:28 <brouberol@cumin1003> START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage [production]
14:27 <brouberol@cumin1003> START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage [production]
14:23 <sukhe> sudo cumin -b2 'A:cp-text' "run-puppet-agent --enable 'merging CR 1290731'": T425441 [production]
14:18 <swfrench-wmf> begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - T428495 [production]
14:16 <brouberol@cumin1003> START - Cookbook sre.hosts.reimage for host an-test-presto1001.eqiad.wmnet with OS bookworm [production]
14:14 <brouberol@cumin1003> START - Cookbook sre.hosts.reimage for host an-test-ui1001.eqiad.wmnet with OS bookworm [production]
14:12 <sukhe> sudo cumin 'A:cp-text' "disable-puppet 'merging CR 1290731'": T425441 [production]
14:11 <swfrench-wmf> restarted navtiming on webperf1003 - T428495 [production]
14:04 <swfrench-wmf> begin rolling restart of confd in drmrs, eqiad, esams, magru - T428495 [production]
14:04 <brouberol@cumin1003> START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm [production]
14:04 <brouberol@cumin1003> START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm [production]
14:02 <brouberol@cumin1003> END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-client1002.eqiad.wmnet with OS bookworm [production]
13:58 <swfrench-wmf> authdns update to direct eqiad-associated etcd clients back to eqiad - T428495 [production]
13:58 <swfrench@dns1004> END - running authdns-update [production]
13:56 <swfrench@dns1004> START - running authdns-update [production]
13:49 <aikochou@deploy1003> helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . [production]
13:44 <aikochou@deploy1003> helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' . [production]
13:31 <bwojtowicz@deploy1003> helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . [production]
13:29 <bwojtowicz@deploy1003> helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' . [production]
13:26 <bwojtowicz@deploy1003> helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . [production]
13:23 <brouberol@cumin1003> END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage [production]
13:22 <bwojtowicz@deploy1003> helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' . [production]