|
2026-09-23
ยง
|
| 22:08 |
<ryankemper> |
[Cirrus] (to be clear the cluster is not serving live traffic, but if I can avoid red I will) |
[production] |
| 22:08 |
<ryankemper> |
[Cirrus] updater still failing in codfw cirrussearch; i've restarted the directly-impacted hosts but not the others. some bulk updates appear to be getting rejected, going to do some targeted restarts and assess impact before considering a broader operation. first up is `cirrussearch2071.codfw.wmnet` which is not the sole holder of any shards therefore should not plunge the cluster into red status |
[production] |
| 22:06 |
<Southparkfan> |
decommission Beta Cluster PHP 8.3 hosts deployment-mediawiki13, deployment-mediawiki14, deployment-jobrunner05, deployment-mwmaint03 - T435393 |
[releng] |
| 21:49 |
<dzahn@cumin2003> |
START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie |
[production] |
| 21:46 |
<Southparkfan> |
cherry-pick https://gerrit.wikimedia.org/r/c/operations/puppet/+/1344336 to Puppetserver - T435393 |
[releng] |
| 21:46 |
<Southparkfan> |
cherry-pick https://gerrit.wikimedia.org/r/c/operations/puppet/+/1344336 to Puppetserver - T435393 |
[deployment-prep] |
| 21:22 |
<rzl@deploy1003> |
Locking from deployment [ALL REPOSITORIES]: incident recovery in progress T439010 |
[production] |
| 21:22 |
<rzl@deploy1003> |
Unlocked for deployment [ALL REPOSITORIES]: incident recovery in progress T439010 (duration: 51m 29s) |
[production] |
| 21:21 |
<cdobbins@cumin1004> |
END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie |
[production] |
| 21:18 |
<Emperor> |
ceph mgr fail on apus-be2005 |
[production] |
| 21:18 |
<Emperor> |
reset-failed then restart ceph-mon on moss-be2003 |
[production] |
| 21:08 |
<marostegui@cumin1004> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2235].codfw.wmnet with reason: needs fixing |
[production] |
| 21:08 |
<marostegui@cumin1004> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2234].codfw.wmnet with reason: needs fixing |
[production] |
| 21:07 |
<marostegui@cumin1004> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2233].codfw.wmnet with reason: needs fixing |
[production] |
| 21:07 |
<marostegui@cumin1004> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2232].codfw.wmnet with reason: needs fixing |
[production] |
| 20:57 |
<ryankemper> |
[Cirrus] cirrussearch codfw back to yellow status. active shard pct = 94.51% |
[production] |
| 20:55 |
<ryankemper> |
[Cirrus] Bump codfw cirrussearch shard recoveries from 5 to 10; cluster not serving live traffic so I'm hoping we have headroom to recover faster |
[production] |
| 20:49 |
<swfrench@dns1004> |
END - running authdns-update |
[production] |
| 20:46 |
<swfrench@dns1004> |
START - running authdns-update |
[production] |
| 20:41 |
<ryankemper> |
[Cirrus] Been restarting all impacted codfw opensearch hosts one at a time (they didn't rejoin the cluster naturally) |
[production] |
| 20:39 |
<cdobbins@cumin1004> |
START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie |
[production] |
| 20:38 |
<cdobbins@cumin1004> |
END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie |
[production] |
| 20:37 |
<hashar> |
gerrit: deleted https://gerrit.wikimedia.org/r/c/mediawiki/core/+/1344362 duplicate Change-Id I67dd7b3c879049a1c98586c365c297eac150b963 of https://gerrit.wikimedia.org/r/c/mediawiki/core/+/1329697 |
[releng] |
| 20:30 |
<rzl@deploy1003> |
Locking from deployment [ALL REPOSITORIES]: incident recovery in progress T439010 |
[production] |
| 20:27 |
<lerickson@deploy1003> |
helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply |
[production] |
| 20:27 |
<lerickson@deploy1003> |
helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply |
[production] |
| 20:26 |
<dancy> |
Upgrading istio to 1.30.5 on gitlab-cloud-runners (staging) (T439035) |
[releng] |
| 20:13 |
<hashar> |
gerrit: reindexing has been completed |
[releng] |
| 20:06 |
<dzahn@dns1004> |
END - running authdns-update |
[production] |
| 20:03 |
<dzahn@dns1004> |
START - running authdns-update |
[production] |
| 19:52 |
<cdobbins@cumin1004> |
START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie |
[production] |
| 19:38 |
<hashar> |
gerrit: started reindexing for accounts, group and projects indices |
[releng] |
| 19:36 |
<hashar> |
gerrit: force online reindexing of the change index over ssh hitting gerrit1003 (switch over). `gerrit index start changes --force` |
[releng] |
| 19:34 |
<sukhe@cumin1004> |
END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp2059.codfw.wmnet with OS trixie |
[production] |
| 19:33 |
<lerickson@deploy1003> |
helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply |
[production] |
| 19:33 |
<volans> |
rebooting arclamp2001.codfw.wmnet |
[production] |
| 19:32 |
<lerickson@deploy1003> |
helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply |
[production] |
| 19:21 |
<andrew@cloudcumin1001> |
END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services |
[admin] |
| 19:20 |
<sukhe@cumin1004> |
DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 979 hosts with reason: power is still coming back on |
[production] |
| 19:18 |
<andrew@cloudcumin1001> |
START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services |
[admin] |
| 19:17 |
<taavi@dns1004> |
END - running authdns-update |
[production] |
| 19:17 |
<andrew@cloudcumin1001> |
END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services |
[admin] |
| 19:17 |
<andrew@cloudcumin1001> |
START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services |
[admin] |
| 19:14 |
<taavi@dns1004> |
START - running authdns-update |
[production] |
| 19:11 |
<wmbot~lucaswerkmeister@tools-bastion-15> |
deployed fa3fbc988f (Toolforge Components Service, push-to-deploy) |
[tools.speedpatrolling] |
| 19:10 |
<taavi@cumin1004> |
END (PASS) - Cookbook sre.gerrit.read-only-toggle (exit_code=0) from gerrit1003.wikimedia.org |
[production] |
| 19:10 |
<taavi@cumin1004> |
START - Cookbook sre.gerrit.read-only-toggle from gerrit1003.wikimedia.org |
[production] |
| 19:10 |
<cdobbins@cumin1004> |
conftool action : set/pooled=yes; selector: name=ncredir6001.* |
[production] |
| 19:08 |
<sukhe@puppetserver1001> |
conftool action : set/pooled=no; selector: dc=codfw,cluster=dnsbox,service=authdns-update |
[production] |
| 18:59 |
<sukhe@cumin1004> |
DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 6:00:00 on 980 hosts with reason: power is still coming back on |
[production] |