|
2026-09-24
§
|
| 05:07 |
<ryankemper@cumin2003> |
END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster search_codfw: Restart codfw following today's power incident to ensure we return to our full expected state - ryankemper@cumin2003 - T439010 |
[production] |
| 01:21 |
<ryankemper@cumin2003> |
START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster search_codfw: Restart codfw following today's power incident to ensure we return to our full expected state - ryankemper@cumin2003 - T439010 |
[production] |
| 01:19 |
<ryankemper> |
[Cirrus] Reverted `node_concurrent_recoveries` to 5 from 10, now that we're back to green |
[production] |
| 01:16 |
<ryankemper> |
[Cirrus] With the restart of `cirrussearch2115`, the codfw cluster has officially reached green status!!! Still working on full verification, but we're almost done here |
[production] |
| 01:14 |
<ryankemper@cumin2003> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2115.codfw.wmnet with reason: Codfw survivor recovery on 2115; temporary chi red expected (T439010) |
[production] |
| 01:11 |
<brett@cumin1004> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on cp2059.codfw.wmnet with reason: failing services but not in service yet |
[production] |
| 01:10 |
<ryankemper@cumin2003> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2109.codfw.wmnet with reason: Codfw survivor recovery on 2109; temporary chi red expected (T439010) |
[production] |
| 01:04 |
<ryankemper@cumin2003> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2104.codfw.wmnet with reason: Codfw survivor recovery on 2104; temporary chi red expected (T439010) |
[production] |
| 01:03 |
<ryankemper> |
[Cirrus] grr, I'd missed some hosts. restarting the last few dangling ones, we're really close to back to green, prob 3-ish more hosts |
[production] |
| 00:40 |
<ryankemper> |
[Cirrus] Great news, we briefly dipped red (same as previous restarts) but went back to yellow almost immediately. AFAICT election went fine, still checking though |
[production] |
| 00:38 |
<ryankemper> |
[Cirrus] Preparing to restart cirrussearch2084 (active cluster manager). With luck, this should restore updater availability (and general cluster green status, after some reshuffling) |
[production] |
| 00:35 |
<ryankemper@cumin2003> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on 55 hosts with reason: Codfw chi elected-manager recovery on 2084; expected brief failover and red state (T439010) |
[production] |
| 00:10 |
<brett@puppetserver1001> |
conftool action : set/pooled=yes; selector: name=cp7011.* |
[production] |
| 00:05 |
<brett@cumin1004> |
END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P{cp7011.magru.wmnet} and A:cp - 7.1.1-2~bpo13+wmf3 () |
[production] |
| 00:00 |
<brett@cumin1004> |
START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P{cp7011.magru.wmnet} and A:cp - 7.1.1-2~bpo13+wmf3 () |
[production] |
|
2026-09-23
§
|
| 23:58 |
<dzahn@cumin2003> |
END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie |
[production] |
| 23:56 |
<brett> |
Switching acme-chief primary from codfw to eqiad - T439010 |
[production] |
| 23:54 |
<brett@cumin1004> |
END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on P{cp7011.magru.wmnet} and A:cp - 7.1.1-2~bpo13+wmf3 () |
[production] |
| 23:49 |
<brett@cumin1004> |
START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P{cp7011.magru.wmnet} and A:cp - 7.1.1-2~bpo13+wmf3 () |
[production] |
| 23:48 |
<brett@cumin1004> |
END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P{cp7001.magru.wmnet} and A:cp - 7.1.1-2~bpo13+wmf3 () |
[production] |
| 23:48 |
<ryankemper> |
[Cirrus] Every host except 2084, which is the current elected chi master, has now been restarted, and shard recoveries healed accordingly. AFAICT we will not be able to revive the updater until we restart this host. Pausing for a few mins to mull things over and get my bearings though, because this restart would be higher-touch than the previous ones |
[production] |
| 23:38 |
<ryankemper@cumin2003> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2108.codfw.wmnet with reason: Codfw survivor recovery on 2108; sequential chi and psi restarts (T439010) |
[production] |
| 23:38 |
<brett@cumin1004> |
START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P{cp7001.magru.wmnet} and A:cp - 7.1.1-2~bpo13+wmf3 () |
[production] |
| 23:35 |
<brett> |
import varnish 7.1.1-2~bpo13+wmf3 into trixie-wikimedia (T438293) |
[production] |
| 23:34 |
<ryankemper@cumin2003> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2107.codfw.wmnet with reason: Codfw survivor recovery on 2107; sequential chi and psi restarts (T439010) |
[production] |
| 23:27 |
<ryankemper@cumin2003> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2085.codfw.wmnet with reason: Codfw survivor recovery on 2085; sequential chi and psi restarts (T439010) |
[production] |
| 23:23 |
<rzl@deploy1003> |
Locking from deployment [ALL REPOSITORIES]: No deployments please, as we're still cleaning up from the codfw power incident T439010. Thursday UTC morning at the earliest, but please ask SRE oncall. |
[production] |
| 23:23 |
<rzl@deploy1003> |
Unlocked for deployment [ALL REPOSITORIES]: incident recovery in progress T439010 (duration: 121m 40s) |
[production] |
| 23:20 |
<ryankemper@cumin2003> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2072.codfw.wmnet with reason: Codfw survivor recovery on 2072; sequential chi and psi restarts (T439010) |
[production] |
| 23:09 |
<ryankemper> |
[Cirrus] rolling cirrussearch2086 next |
[production] |
| 23:08 |
<ryankemper@cumin2003> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2086.codfw.wmnet with reason: Codfw survivor recovery on 2086; sequential chi and omega restarts (T439010) |
[production] |
| 23:01 |
<ryankemper> |
[Cirrus] Doing cirrussearch2114 next |
[production] |
| 22:59 |
<ryankemper@cumin2003> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2114.codfw.wmnet with reason: Codfw survivor recovery on 2114; sequential chi and omega restarts (T439010) |
[production] |
| 22:44 |
<ryankemper@cumin2003> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2106.codfw.wmnet with reason: Codfw chi survivor recovery on 2106; temporary red expected (T439010) |
[production] |
| 22:29 |
<ryankemper> |
[Cirrus] proceeding with manual restart of cirrussearch2105; red status expected, hopefully brief but we'll see |
[production] |
| 22:28 |
<ryankemper@cumin2003> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2105.codfw.wmnet with reason: Codfw chi recovery canary on 2105; temporary service interruption expected (T439010) |
[production] |
| 22:24 |
<ryankemper> |
[Cirrus] s/expected/expect |
[production] |
| 22:23 |
<ryankemper> |
[Cirrus] Alright, I'm getting increasingly convinced that there's no way to restore healthy cluster state without inevitably having to restart sole-shard-holder hosts, which will put the cluster into red status. going to start with just `cirrussearch2105`; I expected red status. silencing alerts first so I don't blow out the channel |
[production] |
| 22:08 |
<ryankemper> |
[Cirrus] (to be clear the cluster is not serving live traffic, but if I can avoid red I will) |
[production] |
| 22:08 |
<ryankemper> |
[Cirrus] updater still failing in codfw cirrussearch; i've restarted the directly-impacted hosts but not the others. some bulk updates appear to be getting rejected, going to do some targeted restarts and assess impact before considering a broader operation. first up is `cirrussearch2071.codfw.wmnet` which is not the sole holder of any shards therefore should not plunge the cluster into red status |
[production] |
| 21:49 |
<dzahn@cumin2003> |
START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie |
[production] |
| 21:22 |
<rzl@deploy1003> |
Locking from deployment [ALL REPOSITORIES]: incident recovery in progress T439010 |
[production] |
| 21:22 |
<rzl@deploy1003> |
Unlocked for deployment [ALL REPOSITORIES]: incident recovery in progress T439010 (duration: 51m 29s) |
[production] |
| 21:21 |
<cdobbins@cumin1004> |
END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie |
[production] |
| 21:18 |
<Emperor> |
ceph mgr fail on apus-be2005 |
[production] |
| 21:18 |
<Emperor> |
reset-failed then restart ceph-mon on moss-be2003 |
[production] |
| 21:08 |
<marostegui@cumin1004> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2235].codfw.wmnet with reason: needs fixing |
[production] |
| 21:08 |
<marostegui@cumin1004> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2234].codfw.wmnet with reason: needs fixing |
[production] |
| 21:07 |
<marostegui@cumin1004> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2233].codfw.wmnet with reason: needs fixing |
[production] |
| 21:07 |
<marostegui@cumin1004> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2232].codfw.wmnet with reason: needs fixing |
[production] |