1-50 of 10000 results (1ms)
2026-09-23 ยง
22:08 <ryankemper> [Cirrus] (to be clear the cluster is not serving live traffic, but if I can avoid red I will) [production]
22:08 <ryankemper> [Cirrus] updater still failing in codfw cirrussearch; i've restarted the directly-impacted hosts but not the others. some bulk updates appear to be getting rejected, going to do some targeted restarts and assess impact before considering a broader operation. first up is `cirrussearch2071.codfw.wmnet` which is not the sole holder of any shards therefore should not plunge the cluster into red status [production]
22:06 <Southparkfan> decommission Beta Cluster PHP 8.3 hosts deployment-mediawiki13, deployment-mediawiki14, deployment-jobrunner05, deployment-mwmaint03 - T435393 [releng]
21:49 <dzahn@cumin2003> START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie [production]
21:46 <Southparkfan> cherry-pick https://gerrit.wikimedia.org/r/c/operations/puppet/+/1344336 to Puppetserver - T435393 [releng]
21:46 <Southparkfan> cherry-pick https://gerrit.wikimedia.org/r/c/operations/puppet/+/1344336 to Puppetserver - T435393 [deployment-prep]
21:22 <rzl@deploy1003> Locking from deployment [ALL REPOSITORIES]: incident recovery in progress T439010 [production]
21:22 <rzl@deploy1003> Unlocked for deployment [ALL REPOSITORIES]: incident recovery in progress T439010 (duration: 51m 29s) [production]
21:21 <cdobbins@cumin1004> END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie [production]
21:18 <Emperor> ceph mgr fail on apus-be2005 [production]
21:18 <Emperor> reset-failed then restart ceph-mon on moss-be2003 [production]
21:08 <marostegui@cumin1004> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2235].codfw.wmnet with reason: needs fixing [production]
21:08 <marostegui@cumin1004> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2234].codfw.wmnet with reason: needs fixing [production]
21:07 <marostegui@cumin1004> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2233].codfw.wmnet with reason: needs fixing [production]
21:07 <marostegui@cumin1004> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2232].codfw.wmnet with reason: needs fixing [production]
20:57 <ryankemper> [Cirrus] cirrussearch codfw back to yellow status. active shard pct = 94.51% [production]
20:55 <ryankemper> [Cirrus] Bump codfw cirrussearch shard recoveries from 5 to 10; cluster not serving live traffic so I'm hoping we have headroom to recover faster [production]
20:49 <swfrench@dns1004> END - running authdns-update [production]
20:46 <swfrench@dns1004> START - running authdns-update [production]
20:41 <ryankemper> [Cirrus] Been restarting all impacted codfw opensearch hosts one at a time (they didn't rejoin the cluster naturally) [production]
20:39 <cdobbins@cumin1004> START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie [production]
20:38 <cdobbins@cumin1004> END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie [production]
20:37 <hashar> gerrit: deleted https://gerrit.wikimedia.org/r/c/mediawiki/core/+/1344362 duplicate Change-Id I67dd7b3c879049a1c98586c365c297eac150b963 of https://gerrit.wikimedia.org/r/c/mediawiki/core/+/1329697 [releng]
20:30 <rzl@deploy1003> Locking from deployment [ALL REPOSITORIES]: incident recovery in progress T439010 [production]
20:27 <lerickson@deploy1003> helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply [production]
20:27 <lerickson@deploy1003> helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply [production]
20:26 <dancy> Upgrading istio to 1.30.5 on gitlab-cloud-runners (staging) (T439035) [releng]
20:13 <hashar> gerrit: reindexing has been completed [releng]
20:06 <dzahn@dns1004> END - running authdns-update [production]
20:03 <dzahn@dns1004> START - running authdns-update [production]
19:52 <cdobbins@cumin1004> START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie [production]
19:38 <hashar> gerrit: started reindexing for accounts, group and projects indices [releng]
19:36 <hashar> gerrit: force online reindexing of the change index over ssh hitting gerrit1003 (switch over). `gerrit index start changes --force` [releng]
19:34 <sukhe@cumin1004> END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp2059.codfw.wmnet with OS trixie [production]
19:33 <lerickson@deploy1003> helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply [production]
19:33 <volans> rebooting arclamp2001.codfw.wmnet [production]
19:32 <lerickson@deploy1003> helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply [production]
19:21 <andrew@cloudcumin1001> END (PASS) - Cookbook wmcs.openstack.restart_openstack (exit_code=0) on deployment codfw1dev for all services [admin]
19:20 <sukhe@cumin1004> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 979 hosts with reason: power is still coming back on [production]
19:18 <andrew@cloudcumin1001> START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services [admin]
19:17 <taavi@dns1004> END - running authdns-update [production]
19:17 <andrew@cloudcumin1001> END (FAIL) - Cookbook wmcs.openstack.restart_openstack (exit_code=99) on deployment codfw1dev for all services [admin]
19:17 <andrew@cloudcumin1001> START - Cookbook wmcs.openstack.restart_openstack on deployment codfw1dev for all services [admin]
19:14 <taavi@dns1004> START - running authdns-update [production]
19:11 <wmbot~lucaswerkmeister@tools-bastion-15> deployed fa3fbc988f (Toolforge Components Service, push-to-deploy) [tools.speedpatrolling]
19:10 <taavi@cumin1004> END (PASS) - Cookbook sre.gerrit.read-only-toggle (exit_code=0) from gerrit1003.wikimedia.org [production]
19:10 <taavi@cumin1004> START - Cookbook sre.gerrit.read-only-toggle from gerrit1003.wikimedia.org [production]
19:10 <cdobbins@cumin1004> conftool action : set/pooled=yes; selector: name=ncredir6001.* [production]
19:08 <sukhe@puppetserver1001> conftool action : set/pooled=no; selector: dc=codfw,cluster=dnsbox,service=authdns-update [production]
18:59 <sukhe@cumin1004> DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 6:00:00 on 980 hosts with reason: power is still coming back on [production]