651-700 of 10000 results (152ms)
2026-07-28 §
07:22 <root@cumin1003> START - Cookbook sre.mysql.pool pool db2154: Maintenance [production]
07:22 <root@cumin1003> START - Cookbook sre.mysql.pool pool db2168: Maintenance [production]
07:16 <cwilliams@cumin1003> dbctl commit (dc=all): 'Depooling db2154 (T431660)', diff saved to https://phabricator.wikimedia.org/P95295 and previous config saved to /var/cache/conftool/dbconfig/20260728-071640-cwilliams.json [production]
07:16 <cwilliams@cumin1003> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2154.codfw.wmnet with reason: Maintenance [production]
07:16 <cwilliams@cumin1003> dbctl commit (dc=all): 'Depooling db2168 (T431660)', diff saved to https://phabricator.wikimedia.org/P95294 and previous config saved to /var/cache/conftool/dbconfig/20260728-071604-cwilliams.json [production]
07:15 <cwilliams@cumin1003> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2168.codfw.wmnet with reason: Maintenance [production]
07:08 <root@cumin1003> START - Cookbook sre.mysql.pool pool db2206: Maintenance [production]
07:02 <cwilliams@cumin1003> dbctl commit (dc=all): 'Depooling db2206 (T431660)', diff saved to https://phabricator.wikimedia.org/P95292 and previous config saved to /var/cache/conftool/dbconfig/20260728-070219-cwilliams.json [production]
07:02 <cwilliams@cumin1003> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2206.codfw.wmnet with reason: Maintenance [production]
06:44 <marostegui> Failover m5 from db1164 to db1228 - T432967 [production]
06:39 <marostegui@cumin1003> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2235].codfw.wmnet,db[1164,1217,1228].eqiad.wmnet with reason: m5 master switch T432967 [production]
04:02 <mwpresync@deploy1003> Pruned MediaWiki: 1.47.0-wmf.10 (duration: 02m 34s) [production]
03:39 <mwpresync@deploy1003> Finished scap sync-world: testwikis to 1.47.0-wmf.13 refs T430832 (duration: 36m 06s) [production]
03:03 <mwpresync@deploy1003> Started scap sync-world: testwikis to 1.47.0-wmf.13 refs T430832 [production]
02:57 <dzahn@cumin1003> END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1004.eqiad.wmnet with OS trixie [production]
02:57 <dzahn@cumin1003> END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - dzahn@cumin1003" [production]
02:55 <dzahn@cumin1003> START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - dzahn@cumin1003" [production]
02:37 <dzahn@cumin1003> END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1004.eqiad.wmnet with reason: host reimage [production]
02:31 <dzahn@cumin1003> START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1004.eqiad.wmnet with reason: host reimage [production]
02:16 <dzahn@cumin1003> START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie [production]
02:15 <dzahn@cumin1003> END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host zuul1004.eqiad.wmnet with OS trixie [production]
01:43 <dzahn@cumin1003> START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie [production]
01:43 <dzahn@cumin1003> END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1004.eqiad.wmnet with OS trixie [production]
01:25 <pt1979@cumin1003> END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-603-eqsin [production]
01:24 <pt1979@cumin1003> START - Cookbook sre.network.tls for network device asw1-603-eqsin [production]
01:12 <pt1979@cumin2003> END (PASS) - Cookbook sre.dns.netbox (exit_code=0) [production]
01:12 <pt1979@cumin2003> END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt for new switches in eqsin - pt1979@cumin2003" [production]
01:12 <pt1979@cumin2003> START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt for new switches in eqsin - pt1979@cumin2003" [production]
01:08 <pt1979@cumin2003> START - Cookbook sre.dns.netbox [production]
00:48 <swfrench@deploy1003> helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply [production]
00:47 <swfrench@deploy1003> helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply [production]
00:47 <swfrench@deploy1003> helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply [production]
00:47 <swfrench@deploy1003> helmfile [codfw] START helmfile.d/services/mw-pretrain: apply [production]
00:26 <mutante> attempting reimage with trixie on zuul1004 re-purposed physical hardware - dcops reported install issue - host was in busybox shell (T427353) [production]
00:24 <dzahn@cumin1003> START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie [production]
2026-07-27 §
23:50 <Amir1> mass deleting vp8 transcodes [production]
23:28 <amastilovic@deploy1003> helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply [production]
23:27 <vriley@cumin1003> END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1004.eqiad.wmnet with OS bullseye [production]
23:26 <amastilovic@deploy1003> helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply [production]
23:25 <amastilovic@deploy1003> helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply [production]
23:25 <amastilovic@deploy1003> helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply [production]
22:39 <maryum> Deploy security fix for T432877 [production]
22:38 <bking@cumin2003> END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1022.eqiad.wmnet with OS bookworm [production]
22:37 <vriley@cumin1003> START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS bullseye [production]
22:32 <sbassett> Deployed security fix for T432789 [production]
22:22 <sbassett> Deployed security patch for T431819 [production]
22:14 <bking@cumin2003> END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1022.eqiad.wmnet with reason: host reimage [production]
22:07 <bking@cumin2003> START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1022.eqiad.wmnet with reason: host reimage [production]
22:01 <RScout-WMF> Deployed security fix for T431819 [production]
22:00 <bking@cumin2003> START - Cookbook sre.hosts.move-vlan for host wdqs2012 [production]