|
2026-09-23
ยง
|
| 21:22 |
<rzl@deploy1003> |
Locking from deployment [ALL REPOSITORIES]: incident recovery in progress T439010 |
[production] |
| 21:22 |
<rzl@deploy1003> |
Unlocked for deployment [ALL REPOSITORIES]: incident recovery in progress T439010 (duration: 51m 29s) |
[production] |
| 21:21 |
<cdobbins@cumin1004> |
END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie |
[production] |
| 21:18 |
<Emperor> |
ceph mgr fail on apus-be2005 |
[production] |
| 21:18 |
<Emperor> |
reset-failed then restart ceph-mon on moss-be2003 |
[production] |
| 21:08 |
<marostegui@cumin1004> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2235].codfw.wmnet with reason: needs fixing |
[production] |
| 21:08 |
<marostegui@cumin1004> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2234].codfw.wmnet with reason: needs fixing |
[production] |
| 21:07 |
<marostegui@cumin1004> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2233].codfw.wmnet with reason: needs fixing |
[production] |
| 21:07 |
<marostegui@cumin1004> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2232].codfw.wmnet with reason: needs fixing |
[production] |
| 20:57 |
<ryankemper> |
[Cirrus] cirrussearch codfw back to yellow status. active shard pct = 94.51% |
[production] |
| 20:55 |
<ryankemper> |
[Cirrus] Bump codfw cirrussearch shard recoveries from 5 to 10; cluster not serving live traffic so I'm hoping we have headroom to recover faster |
[production] |
| 20:49 |
<swfrench@dns1004> |
END - running authdns-update |
[production] |
| 20:46 |
<swfrench@dns1004> |
START - running authdns-update |
[production] |
| 20:41 |
<ryankemper> |
[Cirrus] Been restarting all impacted codfw opensearch hosts one at a time (they didn't rejoin the cluster naturally) |
[production] |
| 20:39 |
<cdobbins@cumin1004> |
START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie |
[production] |
| 20:38 |
<cdobbins@cumin1004> |
END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie |
[production] |
| 20:30 |
<rzl@deploy1003> |
Locking from deployment [ALL REPOSITORIES]: incident recovery in progress T439010 |
[production] |
| 20:27 |
<lerickson@deploy1003> |
helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply |
[production] |
| 20:27 |
<lerickson@deploy1003> |
helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply |
[production] |
| 20:06 |
<dzahn@dns1004> |
END - running authdns-update |
[production] |
| 20:03 |
<dzahn@dns1004> |
START - running authdns-update |
[production] |
| 19:52 |
<cdobbins@cumin1004> |
START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie |
[production] |
| 19:34 |
<sukhe@cumin1004> |
END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp2059.codfw.wmnet with OS trixie |
[production] |
| 19:33 |
<lerickson@deploy1003> |
helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply |
[production] |
| 19:33 |
<volans> |
rebooting arclamp2001.codfw.wmnet |
[production] |
| 19:32 |
<lerickson@deploy1003> |
helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply |
[production] |
| 19:20 |
<sukhe@cumin1004> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 979 hosts with reason: power is still coming back on |
[production] |
| 19:17 |
<taavi@dns1004> |
END - running authdns-update |
[production] |
| 19:14 |
<taavi@dns1004> |
START - running authdns-update |
[production] |
| 19:10 |
<taavi@cumin1004> |
END (PASS) - Cookbook sre.gerrit.read-only-toggle (exit_code=0) from gerrit1003.wikimedia.org |
[production] |
| 19:10 |
<taavi@cumin1004> |
START - Cookbook sre.gerrit.read-only-toggle from gerrit1003.wikimedia.org |
[production] |
| 19:10 |
<cdobbins@cumin1004> |
conftool action : set/pooled=yes; selector: name=ncredir6001.* |
[production] |
| 19:08 |
<sukhe@puppetserver1001> |
conftool action : set/pooled=no; selector: dc=codfw,cluster=dnsbox,service=authdns-update |
[production] |
| 18:59 |
<sukhe@cumin1004> |
DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 6:00:00 on 980 hosts with reason: power is still coming back on |
[production] |
| 18:58 |
<taavi@cumin1004> |
END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) gerrit.discovery.wmnet on all recursors |
[production] |
| 18:58 |
<taavi@cumin1004> |
START - Cookbook sre.dns.wipe-cache gerrit.discovery.wmnet on all recursors |
[production] |
| 18:50 |
<taavi@cumin1004> |
END (PASS) - Cookbook sre.gerrit.localbackup (exit_code=0) Prepare local backup on: gerrit2003.wikimedia.org |
[production] |
| 18:45 |
<sukhe@dns1004> |
END - running authdns-update |
[production] |
| 18:43 |
<sukhe@dns1004> |
START - running authdns-update |
[production] |
| 18:43 |
<taavi@cumin1004> |
START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org |
[production] |
| 18:42 |
<sukhe@puppetserver1001> |
conftool action : set/pooled=yes; selector: dc=codfw,cluster=dnsbox,service=authdns-update |
[production] |
| 18:42 |
<dzahn@cumin2003> |
END (FAIL) - Cookbook sre.gerrit.localbackup (exit_code=99) Prepare local backup on: gerrit2003.wikimedia.org |
[production] |
| 18:42 |
<dzahn@cumin2003> |
START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org |
[production] |
| 18:40 |
<dzahn@cumin2003> |
END (FAIL) - Cookbook sre.gerrit.localbackup (exit_code=99) Prepare local backup on: gerrit2003.wikimedia.org |
[production] |
| 18:40 |
<dzahn@cumin2003> |
START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org |
[production] |
| 18:40 |
<dzahn@cumin2003> |
END (FAIL) - Cookbook sre.gerrit.localbackup (exit_code=99) Prepare local backup on: gerrit2003.wikimedia.org |
[production] |
| 18:40 |
<dzahn@cumin2003> |
START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org |
[production] |
| 18:40 |
<taavi@cumin1004> |
END (PASS) - Cookbook sre.gerrit.localbackup (exit_code=0) Prepare local backup on: gerrit1003.wikimedia.org |
[production] |
| 18:38 |
<cdanis@cumin1004> |
END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) _etcd-client-ssl._tcp.eqsin.wmnet _etcd-client-ssl._tcp.ulsfo.wmnet _etcd-client-ssl._tcp.codfw.wmnet on all recursors |
[production] |
| 18:38 |
<cdanis@cumin1004> |
START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.eqsin.wmnet _etcd-client-ssl._tcp.ulsfo.wmnet _etcd-client-ssl._tcp.codfw.wmnet on all recursors |
[production] |