|
2026-09-24
§
|
| 08:31 |
<btullis@cumin1004> |
START - Cookbook sre.k8s.reboot-nodes rolling reboot on P{dse-k8s-worker1015.eqiad.wmnet} and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad) |
[production] |
| 08:22 |
<XioNoX> |
asw1-b12-drmrs> request system reboot - T437984 |
[production] |
| 08:20 |
<ayounsi@cumin1004> |
END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for drmrs rack B12 |
[production] |
| 08:13 |
<ayounsi@cumin1004> |
START - Cookbook sre.network.depool-rack with action 'depool' for drmrs rack B12 |
[production] |
| 08:06 |
<jelto@cumin1004> |
START - Cookbook sre.gitlab.reboot-runner rolling reboot on A:gitlab-runner |
[production] |
| 08:02 |
<ayounsi@cumin1004> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-b12-drmrs,asw1-b12-drmrs IPv6,asw1-b12-drmrs.mgmt with reason: Switch upgrade |
[production] |
| 07:53 |
<ayounsi@cumin1004> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 20 hosts with reason: Switches upgrade |
[production] |
| 07:52 |
<ayounsi@cumin1004> |
END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool drmrs [reason: switch upgrade, T437984] |
[production] |
| 07:52 |
<ayounsi@cumin1004> |
START - Cookbook sre.dns.admin DNS admin: depool drmrs [reason: switch upgrade, T437984] |
[production] |
| 07:48 |
<jelto@cumin1004> |
END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: version upgrade |
[production] |
| 07:19 |
<jelto@cumin1004> |
START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: version upgrade |
[production] |
| 07:16 |
<jelto@cumin1004> |
END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: version upgrade |
[production] |
| 07:06 |
<jelto@cumin1004> |
START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: version upgrade |
[production] |
| 07:02 |
<jelto@cumin1004> |
END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: version upgrade |
[production] |
| 06:51 |
<jelto@cumin1004> |
START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: version upgrade |
[production] |
| 06:41 |
<kart_> |
staging: Update machinetranslation/MinT to 2026-09-21-112314-production (T437213) |
[production] |
| 06:41 |
<kartik@deploy1003> |
helmfile [staging] DONE helmfile.d/services/machinetranslation: apply |
[production] |
| 06:39 |
<kart_> |
staging: Update machinetranslation/MinT to 2026-09-21-112314-production |
[production] |
| 06:38 |
<kartik@deploy1003> |
helmfile [staging] START helmfile.d/services/machinetranslation: apply |
[production] |
| 06:07 |
<ayounsi@cumin1004> |
END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-e5-codfw |
[production] |
| 06:06 |
<ayounsi@cumin1004> |
START - Cookbook sre.network.tls for network device lsw1-e5-codfw |
[production] |
| 05:07 |
<ryankemper@cumin2003> |
END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster search_codfw: Restart codfw following today's power incident to ensure we return to our full expected state - ryankemper@cumin2003 - T439010 |
[production] |
| 01:21 |
<ryankemper@cumin2003> |
START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster search_codfw: Restart codfw following today's power incident to ensure we return to our full expected state - ryankemper@cumin2003 - T439010 |
[production] |
| 01:19 |
<ryankemper> |
[Cirrus] Reverted `node_concurrent_recoveries` to 5 from 10, now that we're back to green |
[production] |
| 01:16 |
<ryankemper> |
[Cirrus] With the restart of `cirrussearch2115`, the codfw cluster has officially reached green status!!! Still working on full verification, but we're almost done here |
[production] |
| 01:14 |
<ryankemper@cumin2003> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2115.codfw.wmnet with reason: Codfw survivor recovery on 2115; temporary chi red expected (T439010) |
[production] |
| 01:11 |
<brett@cumin1004> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on cp2059.codfw.wmnet with reason: failing services but not in service yet |
[production] |
| 01:10 |
<ryankemper@cumin2003> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2109.codfw.wmnet with reason: Codfw survivor recovery on 2109; temporary chi red expected (T439010) |
[production] |
| 01:04 |
<ryankemper@cumin2003> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2104.codfw.wmnet with reason: Codfw survivor recovery on 2104; temporary chi red expected (T439010) |
[production] |
| 01:03 |
<ryankemper> |
[Cirrus] grr, I'd missed some hosts. restarting the last few dangling ones, we're really close to back to green, prob 3-ish more hosts |
[production] |
| 00:40 |
<ryankemper> |
[Cirrus] Great news, we briefly dipped red (same as previous restarts) but went back to yellow almost immediately. AFAICT election went fine, still checking though |
[production] |
| 00:38 |
<ryankemper> |
[Cirrus] Preparing to restart cirrussearch2084 (active cluster manager). With luck, this should restore updater availability (and general cluster green status, after some reshuffling) |
[production] |
| 00:35 |
<ryankemper@cumin2003> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on 55 hosts with reason: Codfw chi elected-manager recovery on 2084; expected brief failover and red state (T439010) |
[production] |
| 00:10 |
<brett@puppetserver1001> |
conftool action : set/pooled=yes; selector: name=cp7011.* |
[production] |
| 00:05 |
<brett@cumin1004> |
END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P{cp7011.magru.wmnet} and A:cp - 7.1.1-2~bpo13+wmf3 () |
[production] |
| 00:00 |
<brett@cumin1004> |
START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P{cp7011.magru.wmnet} and A:cp - 7.1.1-2~bpo13+wmf3 () |
[production] |
|
2026-09-23
§
|
| 23:58 |
<dzahn@cumin2003> |
END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie |
[production] |
| 23:56 |
<brett> |
Switching acme-chief primary from codfw to eqiad - T439010 |
[production] |
| 23:54 |
<brett@cumin1004> |
END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on P{cp7011.magru.wmnet} and A:cp - 7.1.1-2~bpo13+wmf3 () |
[production] |
| 23:49 |
<brett@cumin1004> |
START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P{cp7011.magru.wmnet} and A:cp - 7.1.1-2~bpo13+wmf3 () |
[production] |
| 23:48 |
<brett@cumin1004> |
END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P{cp7001.magru.wmnet} and A:cp - 7.1.1-2~bpo13+wmf3 () |
[production] |
| 23:48 |
<ryankemper> |
[Cirrus] Every host except 2084, which is the current elected chi master, has now been restarted, and shard recoveries healed accordingly. AFAICT we will not be able to revive the updater until we restart this host. Pausing for a few mins to mull things over and get my bearings though, because this restart would be higher-touch than the previous ones |
[production] |
| 23:38 |
<ryankemper@cumin2003> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2108.codfw.wmnet with reason: Codfw survivor recovery on 2108; sequential chi and psi restarts (T439010) |
[production] |
| 23:38 |
<brett@cumin1004> |
START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P{cp7001.magru.wmnet} and A:cp - 7.1.1-2~bpo13+wmf3 () |
[production] |
| 23:35 |
<brett> |
import varnish 7.1.1-2~bpo13+wmf3 into trixie-wikimedia (T438293) |
[production] |
| 23:34 |
<ryankemper@cumin2003> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2107.codfw.wmnet with reason: Codfw survivor recovery on 2107; sequential chi and psi restarts (T439010) |
[production] |
| 23:27 |
<ryankemper@cumin2003> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2085.codfw.wmnet with reason: Codfw survivor recovery on 2085; sequential chi and psi restarts (T439010) |
[production] |
| 23:23 |
<rzl@deploy1003> |
Locking from deployment [ALL REPOSITORIES]: No deployments please, as we're still cleaning up from the codfw power incident T439010. Thursday UTC morning at the earliest, but please ask SRE oncall. |
[production] |
| 23:23 |
<rzl@deploy1003> |
Unlocked for deployment [ALL REPOSITORIES]: incident recovery in progress T439010 (duration: 121m 40s) |
[production] |
| 23:20 |
<ryankemper@cumin2003> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2072.codfw.wmnet with reason: Codfw survivor recovery on 2072; sequential chi and psi restarts (T439010) |
[production] |