1-50 of 10000 results (54ms)
2026-09-09 §
09:28 <btullis@cumin1003> END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet [production]
09:27 <btullis@cumin1003> START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet [production]
09:03 <brouberol@deploy1003> helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'. [production]
09:02 <brouberol@deploy1003> helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'. [production]
09:01 <cmooney@cumin1004> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 16 hosts with reason: upgrade ssw1-a1-eqiad [production]
08:58 <cmooney@cumin1004> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 22 hosts with reason: upgrade ssw1-a1-eqiad [production]
08:56 <brouberol@deploy1003> helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'. [production]
08:55 <brouberol@deploy1003> helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'. [production]
08:49 <cmooney@cumin1004> END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw [production]
08:49 <cmooney@cumin1004> START - Cookbook sre.network.tls for network device lsw1-d3-codfw [production]
08:48 <cmooney@cumin1004> END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw [production]
08:48 <cmooney@cumin1004> START - Cookbook sre.network.tls for network device ssw1-d1-codfw [production]
08:36 <marostegui@cumin1003> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning sanitarium [production]
08:30 <brouberol@dns1004> END - running authdns-update [production]
08:28 <moritzm> pruned obsolete Bullseye image golang1.15 from the docker registry T416452 [production]
08:28 <brouberol@dns1004> START - running authdns-update [production]
08:06 <marostegui@cumin1003> END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1260: Needs to clone another host from this one [production]
08:06 <marostegui@cumin1003> START - Cookbook sre.mysql.depool depool db1260: Needs to clone another host from this one [production]
08:03 <marostegui@cumin1003> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1260.eqiad.wmnet with reason: Cloning sanitarium [production]
08:01 <btullis@cumin1003> END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet [production]
08:01 <btullis@cumin1003> START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet [production]
08:01 <btullis@cumin1003> END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet [production]
08:00 <btullis@cumin1003> START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet [production]
07:40 <chlod> UTC morning backport window done [production]
07:37 <chlod@deploy1003> Finished scap sync-world: Backport for [[gerrit:1335706|thwikibooks: update tagline and wordmark (T436426)]] (duration: 21m 36s) [production]
07:32 <chlod@deploy1003> chlod, hamishz: Continuing with deployment [production]
07:20 <chlod@deploy1003> chlod, hamishz: Backport for [[gerrit:1335706|thwikibooks: update tagline and wordmark (T436426)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [production]
07:15 <chlod@deploy1003> Started scap sync-world: Backport for [[gerrit:1335706|thwikibooks: update tagline and wordmark (T436426)]] [production]
02:08 <mwpresync@deploy1003> Finished scap build-images: Publishing wmf/next image (duration: 07m 45s) [production]
02:00 <mwpresync@deploy1003> Started scap build-images: Publishing wmf/next image [production]
00:09 <eevans@cumin1003> END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1025.eqiad.wmnet with OS bookworm [production]
2026-09-08 §
23:51 <eevans@cumin1003> END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage [production]
23:48 <eevans@cumin1003> START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage [production]
23:39 <eevans@cumin1003> START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm [production]
23:37 <swfrench@cumin1003> END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1313.eqiad.wmnet [production]
23:37 <swfrench@cumin1003> END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1313.eqiad.wmnet [production]
23:37 <swfrench@cumin1003> START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1313.eqiad.wmnet [production]
23:35 <eevans@cumin1003> END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm [production]
23:30 <eevans@cumin1003> START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm [production]
23:30 <eevans@cumin1003> END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm [production]
23:25 <swfrench@cumin1003> END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1313.eqiad.wmnet with OS trixie [production]
23:19 <eevans@cumin1003> START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm [production]
23:19 <eevans@cumin1003> END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm [production]
23:15 <eevans@cumin1003> START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm [production]
23:14 <eevans@cumin1003> END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART [production]
23:07 <eevans@cumin1003> START - Cookbook sre.hosts.provision for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART [production]
23:07 <eevans@cumin1003> END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm [production]
23:05 <Amir1> dropped 57 tables on db1260 (T437278) [production]
23:03 <swfrench@cumin1003> END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage [production]
23:03 <eevans@cumin1003> START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm [production]