1-50 of 10000 results (56ms)
2026-09-23 ยง
23:09 <ryankemper> [Cirrus] rolling cirrussearch2086 next [production]
23:08 <ryankemper@cumin2003> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2086.codfw.wmnet with reason: Codfw survivor recovery on 2086; sequential chi and omega restarts (T439010) [production]
23:01 <ryankemper> [Cirrus] Doing cirrussearch2114 next [production]
22:59 <ryankemper@cumin2003> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2114.codfw.wmnet with reason: Codfw survivor recovery on 2114; sequential chi and omega restarts (T439010) [production]
22:44 <ryankemper@cumin2003> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2106.codfw.wmnet with reason: Codfw chi survivor recovery on 2106; temporary red expected (T439010) [production]
22:29 <ryankemper> [Cirrus] proceeding with manual restart of cirrussearch2105; red status expected, hopefully brief but we'll see [production]
22:28 <ryankemper@cumin2003> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2105.codfw.wmnet with reason: Codfw chi recovery canary on 2105; temporary service interruption expected (T439010) [production]
22:24 <ryankemper> [Cirrus] s/expected/expect [production]
22:23 <ryankemper> [Cirrus] Alright, I'm getting increasingly convinced that there's no way to restore healthy cluster state without inevitably having to restart sole-shard-holder hosts, which will put the cluster into red status. going to start with just `cirrussearch2105`; I expected red status. silencing alerts first so I don't blow out the channel [production]
22:08 <ryankemper> [Cirrus] (to be clear the cluster is not serving live traffic, but if I can avoid red I will) [production]
22:08 <ryankemper> [Cirrus] updater still failing in codfw cirrussearch; i've restarted the directly-impacted hosts but not the others. some bulk updates appear to be getting rejected, going to do some targeted restarts and assess impact before considering a broader operation. first up is `cirrussearch2071.codfw.wmnet` which is not the sole holder of any shards therefore should not plunge the cluster into red status [production]
21:49 <dzahn@cumin2003> START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie [production]
21:22 <rzl@deploy1003> Locking from deployment [ALL REPOSITORIES]: incident recovery in progress T439010 [production]
21:22 <rzl@deploy1003> Unlocked for deployment [ALL REPOSITORIES]: incident recovery in progress T439010 (duration: 51m 29s) [production]
21:21 <cdobbins@cumin1004> END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie [production]
21:18 <Emperor> ceph mgr fail on apus-be2005 [production]
21:18 <Emperor> reset-failed then restart ceph-mon on moss-be2003 [production]
21:08 <marostegui@cumin1004> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2235].codfw.wmnet with reason: needs fixing [production]
21:08 <marostegui@cumin1004> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2234].codfw.wmnet with reason: needs fixing [production]
21:07 <marostegui@cumin1004> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2233].codfw.wmnet with reason: needs fixing [production]
21:07 <marostegui@cumin1004> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2232].codfw.wmnet with reason: needs fixing [production]
20:57 <ryankemper> [Cirrus] cirrussearch codfw back to yellow status. active shard pct = 94.51% [production]
20:55 <ryankemper> [Cirrus] Bump codfw cirrussearch shard recoveries from 5 to 10; cluster not serving live traffic so I'm hoping we have headroom to recover faster [production]
20:49 <swfrench@dns1004> END - running authdns-update [production]
20:46 <swfrench@dns1004> START - running authdns-update [production]
20:41 <ryankemper> [Cirrus] Been restarting all impacted codfw opensearch hosts one at a time (they didn't rejoin the cluster naturally) [production]
20:39 <cdobbins@cumin1004> START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie [production]
20:38 <cdobbins@cumin1004> END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie [production]
20:30 <rzl@deploy1003> Locking from deployment [ALL REPOSITORIES]: incident recovery in progress T439010 [production]
20:27 <lerickson@deploy1003> helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply [production]
20:27 <lerickson@deploy1003> helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply [production]
20:06 <dzahn@dns1004> END - running authdns-update [production]
20:03 <dzahn@dns1004> START - running authdns-update [production]
19:52 <cdobbins@cumin1004> START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie [production]
19:34 <sukhe@cumin1004> END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp2059.codfw.wmnet with OS trixie [production]
19:33 <lerickson@deploy1003> helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply [production]
19:33 <volans> rebooting arclamp2001.codfw.wmnet [production]
19:32 <lerickson@deploy1003> helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply [production]
19:20 <sukhe@cumin1004> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 979 hosts with reason: power is still coming back on [production]
19:17 <taavi@dns1004> END - running authdns-update [production]
19:14 <taavi@dns1004> START - running authdns-update [production]
19:10 <taavi@cumin1004> END (PASS) - Cookbook sre.gerrit.read-only-toggle (exit_code=0) from gerrit1003.wikimedia.org [production]
19:10 <taavi@cumin1004> START - Cookbook sre.gerrit.read-only-toggle from gerrit1003.wikimedia.org [production]
19:10 <cdobbins@cumin1004> conftool action : set/pooled=yes; selector: name=ncredir6001.* [production]
19:08 <sukhe@puppetserver1001> conftool action : set/pooled=no; selector: dc=codfw,cluster=dnsbox,service=authdns-update [production]
18:59 <sukhe@cumin1004> DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 6:00:00 on 980 hosts with reason: power is still coming back on [production]
18:58 <taavi@cumin1004> END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) gerrit.discovery.wmnet on all recursors [production]
18:58 <taavi@cumin1004> START - Cookbook sre.dns.wipe-cache gerrit.discovery.wmnet on all recursors [production]
18:50 <taavi@cumin1004> END (PASS) - Cookbook sre.gerrit.localbackup (exit_code=0) Prepare local backup on: gerrit2003.wikimedia.org [production]
18:45 <sukhe@dns1004> END - running authdns-update [production]