1-50 of 10000 results (47ms)
2026-09-23 ยง
22:08 <ryankemper> [Cirrus] (to be clear the cluster is not serving live traffic, but if I can avoid red I will) [production]
22:08 <ryankemper> [Cirrus] updater still failing in codfw cirrussearch; i've restarted the directly-impacted hosts but not the others. some bulk updates appear to be getting rejected, going to do some targeted restarts and assess impact before considering a broader operation. first up is `cirrussearch2071.codfw.wmnet` which is not the sole holder of any shards therefore should not plunge the cluster into red status [production]
21:49 <dzahn@cumin2003> START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie [production]
21:22 <rzl@deploy1003> Locking from deployment [ALL REPOSITORIES]: incident recovery in progress T439010 [production]
21:22 <rzl@deploy1003> Unlocked for deployment [ALL REPOSITORIES]: incident recovery in progress T439010 (duration: 51m 29s) [production]
21:21 <cdobbins@cumin1004> END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie [production]
21:18 <Emperor> ceph mgr fail on apus-be2005 [production]
21:18 <Emperor> reset-failed then restart ceph-mon on moss-be2003 [production]
21:08 <marostegui@cumin1004> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2235].codfw.wmnet with reason: needs fixing [production]
21:08 <marostegui@cumin1004> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2234].codfw.wmnet with reason: needs fixing [production]
21:07 <marostegui@cumin1004> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2233].codfw.wmnet with reason: needs fixing [production]
21:07 <marostegui@cumin1004> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2232].codfw.wmnet with reason: needs fixing [production]
20:57 <ryankemper> [Cirrus] cirrussearch codfw back to yellow status. active shard pct = 94.51% [production]
20:55 <ryankemper> [Cirrus] Bump codfw cirrussearch shard recoveries from 5 to 10; cluster not serving live traffic so I'm hoping we have headroom to recover faster [production]
20:49 <swfrench@dns1004> END - running authdns-update [production]
20:46 <swfrench@dns1004> START - running authdns-update [production]
20:41 <ryankemper> [Cirrus] Been restarting all impacted codfw opensearch hosts one at a time (they didn't rejoin the cluster naturally) [production]
20:39 <cdobbins@cumin1004> START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie [production]
20:38 <cdobbins@cumin1004> END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie [production]
20:30 <rzl@deploy1003> Locking from deployment [ALL REPOSITORIES]: incident recovery in progress T439010 [production]
20:27 <lerickson@deploy1003> helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply [production]
20:27 <lerickson@deploy1003> helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply [production]
20:06 <dzahn@dns1004> END - running authdns-update [production]
20:03 <dzahn@dns1004> START - running authdns-update [production]
19:52 <cdobbins@cumin1004> START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie [production]
19:34 <sukhe@cumin1004> END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp2059.codfw.wmnet with OS trixie [production]
19:33 <lerickson@deploy1003> helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply [production]
19:33 <volans> rebooting arclamp2001.codfw.wmnet [production]
19:32 <lerickson@deploy1003> helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply [production]
19:20 <sukhe@cumin1004> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 979 hosts with reason: power is still coming back on [production]
19:17 <taavi@dns1004> END - running authdns-update [production]
19:14 <taavi@dns1004> START - running authdns-update [production]
19:10 <taavi@cumin1004> END (PASS) - Cookbook sre.gerrit.read-only-toggle (exit_code=0) from gerrit1003.wikimedia.org [production]
19:10 <taavi@cumin1004> START - Cookbook sre.gerrit.read-only-toggle from gerrit1003.wikimedia.org [production]
19:10 <cdobbins@cumin1004> conftool action : set/pooled=yes; selector: name=ncredir6001.* [production]
19:08 <sukhe@puppetserver1001> conftool action : set/pooled=no; selector: dc=codfw,cluster=dnsbox,service=authdns-update [production]
18:59 <sukhe@cumin1004> DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 6:00:00 on 980 hosts with reason: power is still coming back on [production]
18:58 <taavi@cumin1004> END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) gerrit.discovery.wmnet on all recursors [production]
18:58 <taavi@cumin1004> START - Cookbook sre.dns.wipe-cache gerrit.discovery.wmnet on all recursors [production]
18:50 <taavi@cumin1004> END (PASS) - Cookbook sre.gerrit.localbackup (exit_code=0) Prepare local backup on: gerrit2003.wikimedia.org [production]
18:45 <sukhe@dns1004> END - running authdns-update [production]
18:43 <sukhe@dns1004> START - running authdns-update [production]
18:43 <taavi@cumin1004> START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org [production]
18:42 <sukhe@puppetserver1001> conftool action : set/pooled=yes; selector: dc=codfw,cluster=dnsbox,service=authdns-update [production]
18:42 <dzahn@cumin2003> END (FAIL) - Cookbook sre.gerrit.localbackup (exit_code=99) Prepare local backup on: gerrit2003.wikimedia.org [production]
18:42 <dzahn@cumin2003> START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org [production]
18:40 <dzahn@cumin2003> END (FAIL) - Cookbook sre.gerrit.localbackup (exit_code=99) Prepare local backup on: gerrit2003.wikimedia.org [production]
18:40 <dzahn@cumin2003> START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org [production]
18:40 <dzahn@cumin2003> END (FAIL) - Cookbook sre.gerrit.localbackup (exit_code=99) Prepare local backup on: gerrit2003.wikimedia.org [production]
18:40 <dzahn@cumin2003> START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org [production]