|
2026-09-23
§
|
| 23:58 |
<dzahn@cumin2003> |
END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie |
[production] |
| 23:58 |
<eileen> |
civicrm upgraded from 4fc5d217 to 11d9e123 |
[fundraising] |
| 23:56 |
<brett> |
Switching acme-chief primary from codfw to eqiad - T439010 |
[production] |
| 23:54 |
<brett@cumin1004> |
END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on P{cp7011.magru.wmnet} and A:cp - 7.1.1-2~bpo13+wmf3 () |
[production] |
| 23:49 |
<brett@cumin1004> |
START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P{cp7011.magru.wmnet} and A:cp - 7.1.1-2~bpo13+wmf3 () |
[production] |
| 23:48 |
<brett@cumin1004> |
END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P{cp7001.magru.wmnet} and A:cp - 7.1.1-2~bpo13+wmf3 () |
[production] |
| 23:48 |
<ryankemper> |
[Cirrus] Every host except 2084, which is the current elected chi master, has now been restarted, and shard recoveries healed accordingly. AFAICT we will not be able to revive the updater until we restart this host. Pausing for a few mins to mull things over and get my bearings though, because this restart would be higher-touch than the previous ones |
[production] |
| 23:38 |
<ryankemper@cumin2003> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2108.codfw.wmnet with reason: Codfw survivor recovery on 2108; sequential chi and psi restarts (T439010) |
[production] |
| 23:38 |
<brett@cumin1004> |
START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P{cp7001.magru.wmnet} and A:cp - 7.1.1-2~bpo13+wmf3 () |
[production] |
| 23:35 |
<brett> |
import varnish 7.1.1-2~bpo13+wmf3 into trixie-wikimedia (T438293) |
[production] |
| 23:34 |
<ryankemper@cumin2003> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2107.codfw.wmnet with reason: Codfw survivor recovery on 2107; sequential chi and psi restarts (T439010) |
[production] |
| 23:27 |
<ryankemper@cumin2003> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2085.codfw.wmnet with reason: Codfw survivor recovery on 2085; sequential chi and psi restarts (T439010) |
[production] |
| 23:23 |
<rzl@deploy1003> |
Locking from deployment [ALL REPOSITORIES]: No deployments please, as we're still cleaning up from the codfw power incident T439010. Thursday UTC morning at the earliest, but please ask SRE oncall. |
[production] |
| 23:23 |
<rzl@deploy1003> |
Unlocked for deployment [ALL REPOSITORIES]: incident recovery in progress T439010 (duration: 121m 40s) |
[production] |
| 23:20 |
<ryankemper@cumin2003> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2072.codfw.wmnet with reason: Codfw survivor recovery on 2072; sequential chi and psi restarts (T439010) |
[production] |
| 23:09 |
<ryankemper> |
[Cirrus] rolling cirrussearch2086 next |
[production] |
| 23:08 |
<ryankemper@cumin2003> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2086.codfw.wmnet with reason: Codfw survivor recovery on 2086; sequential chi and omega restarts (T439010) |
[production] |
| 23:01 |
<ryankemper> |
[Cirrus] Doing cirrussearch2114 next |
[production] |
| 22:59 |
<ryankemper@cumin2003> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2114.codfw.wmnet with reason: Codfw survivor recovery on 2114; sequential chi and omega restarts (T439010) |
[production] |
| 22:44 |
<ryankemper@cumin2003> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2106.codfw.wmnet with reason: Codfw chi survivor recovery on 2106; temporary red expected (T439010) |
[production] |
| 22:29 |
<ryankemper> |
[Cirrus] proceeding with manual restart of cirrussearch2105; red status expected, hopefully brief but we'll see |
[production] |
| 22:28 |
<ryankemper@cumin2003> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2105.codfw.wmnet with reason: Codfw chi recovery canary on 2105; temporary service interruption expected (T439010) |
[production] |
| 22:24 |
<ryankemper> |
[Cirrus] s/expected/expect |
[production] |
| 22:23 |
<ryankemper> |
[Cirrus] Alright, I'm getting increasingly convinced that there's no way to restore healthy cluster state without inevitably having to restart sole-shard-holder hosts, which will put the cluster into red status. going to start with just `cirrussearch2105`; I expected red status. silencing alerts first so I don't blow out the channel |
[production] |
| 22:08 |
<ryankemper> |
[Cirrus] (to be clear the cluster is not serving live traffic, but if I can avoid red I will) |
[production] |
| 22:08 |
<ryankemper> |
[Cirrus] updater still failing in codfw cirrussearch; i've restarted the directly-impacted hosts but not the others. some bulk updates appear to be getting rejected, going to do some targeted restarts and assess impact before considering a broader operation. first up is `cirrussearch2071.codfw.wmnet` which is not the sole holder of any shards therefore should not plunge the cluster into red status |
[production] |
| 22:06 |
<Southparkfan> |
decommission Beta Cluster PHP 8.3 hosts deployment-mediawiki13, deployment-mediawiki14, deployment-jobrunner05, deployment-mwmaint03 - T435393 |
[releng] |
| 21:49 |
<dzahn@cumin2003> |
START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie |
[production] |
| 21:46 |
<Southparkfan> |
cherry-pick https://gerrit.wikimedia.org/r/c/operations/puppet/+/1344336 to Puppetserver - T435393 |
[deployment-prep] |
| 21:46 |
<Southparkfan> |
cherry-pick https://gerrit.wikimedia.org/r/c/operations/puppet/+/1344336 to Puppetserver - T435393 |
[releng] |
| 21:22 |
<rzl@deploy1003> |
Locking from deployment [ALL REPOSITORIES]: incident recovery in progress T439010 |
[production] |
| 21:22 |
<rzl@deploy1003> |
Unlocked for deployment [ALL REPOSITORIES]: incident recovery in progress T439010 (duration: 51m 29s) |
[production] |
| 21:21 |
<cdobbins@cumin1004> |
END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie |
[production] |
| 21:18 |
<Emperor> |
ceph mgr fail on apus-be2005 |
[production] |
| 21:18 |
<Emperor> |
reset-failed then restart ceph-mon on moss-be2003 |
[production] |
| 21:08 |
<marostegui@cumin1004> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2235].codfw.wmnet with reason: needs fixing |
[production] |
| 21:08 |
<marostegui@cumin1004> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2234].codfw.wmnet with reason: needs fixing |
[production] |
| 21:07 |
<marostegui@cumin1004> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2233].codfw.wmnet with reason: needs fixing |
[production] |
| 21:07 |
<marostegui@cumin1004> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2232].codfw.wmnet with reason: needs fixing |
[production] |
| 20:57 |
<ryankemper> |
[Cirrus] cirrussearch codfw back to yellow status. active shard pct = 94.51% |
[production] |
| 20:55 |
<ryankemper> |
[Cirrus] Bump codfw cirrussearch shard recoveries from 5 to 10; cluster not serving live traffic so I'm hoping we have headroom to recover faster |
[production] |
| 20:49 |
<swfrench@dns1004> |
END - running authdns-update |
[production] |
| 20:46 |
<swfrench@dns1004> |
START - running authdns-update |
[production] |
| 20:41 |
<ryankemper> |
[Cirrus] Been restarting all impacted codfw opensearch hosts one at a time (they didn't rejoin the cluster naturally) |
[production] |
| 20:39 |
<cdobbins@cumin1004> |
START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie |
[production] |
| 20:38 |
<cdobbins@cumin1004> |
END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie |
[production] |
| 20:37 |
<hashar> |
gerrit: deleted https://gerrit.wikimedia.org/r/c/mediawiki/core/+/1344362 duplicate Change-Id I67dd7b3c879049a1c98586c365c297eac150b963 of https://gerrit.wikimedia.org/r/c/mediawiki/core/+/1329697 |
[releng] |