101-150 of 10000 results (18ms)
2026-09-24 ยง
20:25 <brett@cumin1004> END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on A:cp-text_ulsfo - 7.1.1-2~bpo13+wmf3 () [production]
20:25 <brett@cumin1004> END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on A:cp-upload_ulsfo - 7.1.1-2~bpo13+wmf3 () [production]
20:19 <kemayo@deploy1003> Finished scap sync-world: Backport for [[gerrit:1344714|EditCheck: add some statsv tracking of check/suggestion actions (T438916)]] (duration: 11m 23s) [production]
20:14 <kemayo@deploy1003> kemayo: Continuing with deployment [production]
20:12 <kemayo@deploy1003> kemayo: Backport for [[gerrit:1344714|EditCheck: add some statsv tracking of check/suggestion actions (T438916)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [production]
20:09 <btullis@cumin1004> START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet [production]
20:09 <btullis@cumin1004> END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet [production]
20:09 <btullis@cumin1004> START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet [production]
20:08 <kemayo@deploy1003> Started scap sync-world: Backport for [[gerrit:1344714|EditCheck: add some statsv tracking of check/suggestion actions (T438916)]] [production]
20:01 <btullis@cumin1004> END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet [production]
19:57 <cjming@deploy1003> helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply [production]
19:56 <cjming@deploy1003> helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply [production]
19:56 <cdobbins@cumin1004> END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie [production]
19:50 <brett@cumin1004> END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on A:cp-upload_magru and not P{cp7011.magru.wmnet} and A:cp - 7.1.1-2~bpo13+wmf3 () [production]
19:46 <vriley@cumin1004> START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie [production]
19:36 <ryankemper> [Cirrus] All cirrus pools are serving again. Actively monitoring while the system returns to equilibrium, but all initial indications are that things are as they should be [production]
19:34 <ebernhardson@deploy1003> helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply [production]
19:34 <ebernhardson@deploy1003> helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply [production]
19:33 <ryankemper@cumin2003> END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) pool search-omega in codfw: maintenance [production]
19:31 <btullis@cumin1004> START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet [production]
19:31 <btullis@cumin1004> END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet [production]
19:31 <btullis@cumin1004> START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet [production]
19:29 <ebernhardson@deploy1003> helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply [production]
19:29 <ebernhardson@deploy1003> helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply [production]
19:28 <ryankemper@cumin2003> START - Cookbook sre.discovery.service-route pool search-omega in codfw: maintenance [production]
19:27 <cdanis@cumin1004> conftool action : set/pooled=true; selector: name=codfw,dnsdisc=k8s-ingress-aux-ro [production]
19:26 <ryankemper> [Cirrus] nevermind, that's just the cookbook assuming the DNS record should exist, which it doesn't because chi/psi/omega all share `search.svc.$DC.wmnet` [production]
19:25 <btullis@cumin1004> END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1012.eqiad.wmnet [production]
19:24 <ryankemper> [Cirrus] `dns.resolver.NoAnswer: The DNS response does not contain an answer to the question: search-psi.svc.eqiad.wmnet` checking briefly if this is real failure or just some TTL wonkiness [production]
19:23 <ryankemper@cumin2003> END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) pool search-psi in codfw: maintenance [production]
19:20 <ebernhardson@deploy1003> helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply [production]
19:20 <ebernhardson@deploy1003> helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply [production]
19:18 <dzahn@cumin2003> END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie [production]
19:18 <ryankemper@cumin2003> START - Cookbook sre.discovery.service-route pool search-psi in codfw: maintenance [production]
19:17 <ryankemper> [Cirrus] codfw chi (big cluster) repooled; metrics are already improving, I see poolcounter rejections dropping significantly [production]
19:17 <ryankemper@cumin2003> END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool search in codfw: maintenance [production]
19:17 <cdanis@cumin1004> conftool action : set/ttl=300; selector: name=codfw [production]
19:13 <cdobbins@cumin1004> START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie [production]
19:12 <ryankemper@cumin2003> START - Cookbook sre.discovery.service-route pool search in codfw: maintenance [production]
19:11 <ryankemper> [Cirrus] Repooling codfw, chi first followed by the small clusters [production]
19:11 <cdanis@cumin1004> conftool action : set/pooled=true; selector: name=codfw,dnsdisc=(kartotherian|tegola-vector-tiles) [production]
19:07 <cdobbins@cumin1004> END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie [production]
19:02 <jhuneidi@deploy1003> Finished scap sync-world: Backport for [[gerrit:1344753|REST: restore PageContentHelper::checkAccess (fix live breakage)]] (duration: 10m 15s) [production]
18:57 <jhuneidi@deploy1003> daniel, jhuneidi: Continuing with deployment [production]
18:56 <jhuneidi@deploy1003> daniel, jhuneidi: Backport for [[gerrit:1344753|REST: restore PageContentHelper::checkAccess (fix live breakage)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [production]
18:55 <btullis@cumin1004> START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet [production]
18:55 <btullis@cumin1004> END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet [production]
18:55 <btullis@cumin1004> START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet [production]
18:52 <jhuneidi@deploy1003> Started scap sync-world: Backport for [[gerrit:1344753|REST: restore PageContentHelper::checkAccess (fix live breakage)]] [production]
18:49 <ryankemper@cumin2003> END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-internal-scholarly in codfw: maintenance [production]