851-900 of 10000 results (8ms)
2026-09-23 ยง
22:59 <ryankemper@cumin2003> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2114.codfw.wmnet with reason: Codfw survivor recovery on 2114; sequential chi and omega restarts (T439010) [production]
22:44 <ryankemper@cumin2003> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2106.codfw.wmnet with reason: Codfw chi survivor recovery on 2106; temporary red expected (T439010) [production]
22:29 <ryankemper> [Cirrus] proceeding with manual restart of cirrussearch2105; red status expected, hopefully brief but we'll see [production]
22:28 <ryankemper@cumin2003> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2105.codfw.wmnet with reason: Codfw chi recovery canary on 2105; temporary service interruption expected (T439010) [production]
22:24 <ryankemper> [Cirrus] s/expected/expect [production]
22:23 <ryankemper> [Cirrus] Alright, I'm getting increasingly convinced that there's no way to restore healthy cluster state without inevitably having to restart sole-shard-holder hosts, which will put the cluster into red status. going to start with just `cirrussearch2105`; I expected red status. silencing alerts first so I don't blow out the channel [production]
22:08 <ryankemper> [Cirrus] (to be clear the cluster is not serving live traffic, but if I can avoid red I will) [production]
22:08 <ryankemper> [Cirrus] updater still failing in codfw cirrussearch; i've restarted the directly-impacted hosts but not the others. some bulk updates appear to be getting rejected, going to do some targeted restarts and assess impact before considering a broader operation. first up is `cirrussearch2071.codfw.wmnet` which is not the sole holder of any shards therefore should not plunge the cluster into red status [production]
22:06 <Southparkfan> decommission Beta Cluster PHP 8.3 hosts deployment-mediawiki13, deployment-mediawiki14, deployment-jobrunner05, deployment-mwmaint03 - T435393 [releng]
21:49 <dzahn@cumin2003> START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie [production]
21:46 <Southparkfan> cherry-pick https://gerrit.wikimedia.org/r/c/operations/puppet/+/1344336 to Puppetserver - T435393 [releng]
21:46 <Southparkfan> cherry-pick https://gerrit.wikimedia.org/r/c/operations/puppet/+/1344336 to Puppetserver - T435393 [deployment-prep]
21:22 <rzl@deploy1003> Locking from deployment [ALL REPOSITORIES]: incident recovery in progress T439010 [production]
21:22 <rzl@deploy1003> Unlocked for deployment [ALL REPOSITORIES]: incident recovery in progress T439010 (duration: 51m 29s) [production]
21:21 <cdobbins@cumin1004> END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie [production]
21:18 <Emperor> ceph mgr fail on apus-be2005 [production]
21:18 <Emperor> reset-failed then restart ceph-mon on moss-be2003 [production]
21:08 <marostegui@cumin1004> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2235].codfw.wmnet with reason: needs fixing [production]
21:08 <marostegui@cumin1004> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2234].codfw.wmnet with reason: needs fixing [production]
21:07 <marostegui@cumin1004> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2233].codfw.wmnet with reason: needs fixing [production]
21:07 <marostegui@cumin1004> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2232].codfw.wmnet with reason: needs fixing [production]
20:57 <ryankemper> [Cirrus] cirrussearch codfw back to yellow status. active shard pct = 94.51% [production]
20:55 <ryankemper> [Cirrus] Bump codfw cirrussearch shard recoveries from 5 to 10; cluster not serving live traffic so I'm hoping we have headroom to recover faster [production]
20:49 <swfrench@dns1004> END - running authdns-update [production]
20:46 <swfrench@dns1004> START - running authdns-update [production]
20:41 <ryankemper> [Cirrus] Been restarting all impacted codfw opensearch hosts one at a time (they didn't rejoin the cluster naturally) [production]
20:39 <cdobbins@cumin1004> START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie [production]
20:38 <cdobbins@cumin1004> END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie [production]
20:37 <hashar> gerrit: deleted https://gerrit.wikimedia.org/r/c/mediawiki/core/+/1344362 duplicate Change-Id I67dd7b3c879049a1c98586c365c297eac150b963 of https://gerrit.wikimedia.org/r/c/mediawiki/core/+/1329697 [releng]
20:30 <rzl@deploy1003> Locking from deployment [ALL REPOSITORIES]: incident recovery in progress T439010 [production]
20:27 <lerickson@deploy1003> helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply [production]
20:27 <lerickson@deploy1003> helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply [production]
20:26 <dancy> Upgrading istio to 1.30.5 on gitlab-cloud-runners (staging) (T439035) [releng]
20:13 <hashar> gerrit: reindexing has been completed [releng]
20:06 <dzahn@dns1004> END - running authdns-update [production]
20:03 <dzahn@dns1004> START - running authdns-update [production]
19:52 <cdobbins@cumin1004> START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie [production]
19:38 <hashar> gerrit: started reindexing for accounts, group and projects indices [releng]
19:36 <hashar> gerrit: force online reindexing of the change index over ssh hitting gerrit1003 (switch over). `gerrit index start changes --force` [releng]
19:34 <sukhe@cumin1004> END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp2059.codfw.wmnet with OS trixie [production]
19:33 <lerickson@deploy1003> helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply [production]
19:33 <volans> rebooting arclamp2001.codfw.wmnet [production]
19:32 <lerickson@deploy1003> helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply [production]
19:21 <andrew@cloudcumin1001> END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services [admin]
19:20 <sukhe@cumin1004> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 979 hosts with reason: power is still coming back on [production]
19:18 <andrew@cloudcumin1001> START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services [admin]
19:17 <taavi@dns1004> END - running authdns-update [production]
19:17 <andrew@cloudcumin1001> END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services [admin]
19:17 <andrew@cloudcumin1001> START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services [admin]
19:14 <taavi@dns1004> START - running authdns-update [production]