|
2026-09-23
ยง
|
| 23:09 |
<ryankemper> |
[Cirrus] rolling cirrussearch2086 next |
[production] |
| 23:08 |
<ryankemper@cumin2003> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2086.codfw.wmnet with reason: Codfw survivor recovery on 2086; sequential chi and omega restarts (T439010) |
[production] |
| 23:01 |
<ryankemper> |
[Cirrus] Doing cirrussearch2114 next |
[production] |
| 22:59 |
<ryankemper@cumin2003> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2114.codfw.wmnet with reason: Codfw survivor recovery on 2114; sequential chi and omega restarts (T439010) |
[production] |
| 22:44 |
<ryankemper@cumin2003> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2106.codfw.wmnet with reason: Codfw chi survivor recovery on 2106; temporary red expected (T439010) |
[production] |
| 22:29 |
<ryankemper> |
[Cirrus] proceeding with manual restart of cirrussearch2105; red status expected, hopefully brief but we'll see |
[production] |
| 22:28 |
<ryankemper@cumin2003> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2105.codfw.wmnet with reason: Codfw chi recovery canary on 2105; temporary service interruption expected (T439010) |
[production] |
| 22:24 |
<ryankemper> |
[Cirrus] s/expected/expect |
[production] |
| 22:23 |
<ryankemper> |
[Cirrus] Alright, I'm getting increasingly convinced that there's no way to restore healthy cluster state without inevitably having to restart sole-shard-holder hosts, which will put the cluster into red status. going to start with just `cirrussearch2105`; I expected red status. silencing alerts first so I don't blow out the channel |
[production] |
| 22:08 |
<ryankemper> |
[Cirrus] (to be clear the cluster is not serving live traffic, but if I can avoid red I will) |
[production] |
| 22:08 |
<ryankemper> |
[Cirrus] updater still failing in codfw cirrussearch; i've restarted the directly-impacted hosts but not the others. some bulk updates appear to be getting rejected, going to do some targeted restarts and assess impact before considering a broader operation. first up is `cirrussearch2071.codfw.wmnet` which is not the sole holder of any shards therefore should not plunge the cluster into red status |
[production] |
| 21:49 |
<dzahn@cumin2003> |
START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie |
[production] |
| 21:22 |
<rzl@deploy1003> |
Locking from deployment [ALL REPOSITORIES]: incident recovery in progress T439010 |
[production] |
| 21:22 |
<rzl@deploy1003> |
Unlocked for deployment [ALL REPOSITORIES]: incident recovery in progress T439010 (duration: 51m 29s) |
[production] |
| 21:21 |
<cdobbins@cumin1004> |
END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie |
[production] |
| 21:18 |
<Emperor> |
ceph mgr fail on apus-be2005 |
[production] |
| 21:18 |
<Emperor> |
reset-failed then restart ceph-mon on moss-be2003 |
[production] |
| 21:08 |
<marostegui@cumin1004> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2235].codfw.wmnet with reason: needs fixing |
[production] |
| 21:08 |
<marostegui@cumin1004> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2234].codfw.wmnet with reason: needs fixing |
[production] |
| 21:07 |
<marostegui@cumin1004> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2233].codfw.wmnet with reason: needs fixing |
[production] |
| 21:07 |
<marostegui@cumin1004> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2232].codfw.wmnet with reason: needs fixing |
[production] |
| 20:57 |
<ryankemper> |
[Cirrus] cirrussearch codfw back to yellow status. active shard pct = 94.51% |
[production] |
| 20:55 |
<ryankemper> |
[Cirrus] Bump codfw cirrussearch shard recoveries from 5 to 10; cluster not serving live traffic so I'm hoping we have headroom to recover faster |
[production] |
| 20:49 |
<swfrench@dns1004> |
END - running authdns-update |
[production] |
| 20:46 |
<swfrench@dns1004> |
START - running authdns-update |
[production] |
| 20:41 |
<ryankemper> |
[Cirrus] Been restarting all impacted codfw opensearch hosts one at a time (they didn't rejoin the cluster naturally) |
[production] |
| 20:39 |
<cdobbins@cumin1004> |
START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie |
[production] |
| 20:38 |
<cdobbins@cumin1004> |
END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie |
[production] |
| 20:30 |
<rzl@deploy1003> |
Locking from deployment [ALL REPOSITORIES]: incident recovery in progress T439010 |
[production] |
| 20:27 |
<lerickson@deploy1003> |
helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply |
[production] |
| 20:27 |
<lerickson@deploy1003> |
helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply |
[production] |
| 20:06 |
<dzahn@dns1004> |
END - running authdns-update |
[production] |
| 20:03 |
<dzahn@dns1004> |
START - running authdns-update |
[production] |
| 19:52 |
<cdobbins@cumin1004> |
START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie |
[production] |
| 19:34 |
<sukhe@cumin1004> |
END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp2059.codfw.wmnet with OS trixie |
[production] |
| 19:33 |
<lerickson@deploy1003> |
helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply |
[production] |
| 19:33 |
<volans> |
rebooting arclamp2001.codfw.wmnet |
[production] |
| 19:32 |
<lerickson@deploy1003> |
helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply |
[production] |
| 19:20 |
<sukhe@cumin1004> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 979 hosts with reason: power is still coming back on |
[production] |
| 19:17 |
<taavi@dns1004> |
END - running authdns-update |
[production] |
| 19:14 |
<taavi@dns1004> |
START - running authdns-update |
[production] |
| 19:10 |
<taavi@cumin1004> |
END (PASS) - Cookbook sre.gerrit.read-only-toggle (exit_code=0) from gerrit1003.wikimedia.org |
[production] |
| 19:10 |
<taavi@cumin1004> |
START - Cookbook sre.gerrit.read-only-toggle from gerrit1003.wikimedia.org |
[production] |
| 19:10 |
<cdobbins@cumin1004> |
conftool action : set/pooled=yes; selector: name=ncredir6001.* |
[production] |
| 19:08 |
<sukhe@puppetserver1001> |
conftool action : set/pooled=no; selector: dc=codfw,cluster=dnsbox,service=authdns-update |
[production] |
| 18:59 |
<sukhe@cumin1004> |
DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 6:00:00 on 980 hosts with reason: power is still coming back on |
[production] |
| 18:58 |
<taavi@cumin1004> |
END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) gerrit.discovery.wmnet on all recursors |
[production] |
| 18:58 |
<taavi@cumin1004> |
START - Cookbook sre.dns.wipe-cache gerrit.discovery.wmnet on all recursors |
[production] |
| 18:50 |
<taavi@cumin1004> |
END (PASS) - Cookbook sre.gerrit.localbackup (exit_code=0) Prepare local backup on: gerrit2003.wikimedia.org |
[production] |
| 18:45 |
<sukhe@dns1004> |
END - running authdns-update |
[production] |