|
2026-09-24
ยง
|
| 20:34 |
<ebernhardson@deploy1003> |
helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply |
[production] |
| 20:25 |
<brett@cumin1004> |
END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on A:cp-text_ulsfo - 7.1.1-2~bpo13+wmf3 () |
[production] |
| 20:25 |
<brett@cumin1004> |
END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on A:cp-upload_ulsfo - 7.1.1-2~bpo13+wmf3 () |
[production] |
| 20:19 |
<kemayo@deploy1003> |
Finished scap sync-world: Backport for [[gerrit:1344714|EditCheck: add some statsv tracking of check/suggestion actions (T438916)]] (duration: 11m 23s) |
[production] |
| 20:14 |
<kemayo@deploy1003> |
kemayo: Continuing with deployment |
[production] |
| 20:12 |
<kemayo@deploy1003> |
kemayo: Backport for [[gerrit:1344714|EditCheck: add some statsv tracking of check/suggestion actions (T438916)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. |
[production] |
| 20:09 |
<btullis@cumin1004> |
START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet |
[production] |
| 20:09 |
<btullis@cumin1004> |
END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet |
[production] |
| 20:09 |
<btullis@cumin1004> |
START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet |
[production] |
| 20:08 |
<kemayo@deploy1003> |
Started scap sync-world: Backport for [[gerrit:1344714|EditCheck: add some statsv tracking of check/suggestion actions (T438916)]] |
[production] |
| 20:01 |
<btullis@cumin1004> |
END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet |
[production] |
| 19:57 |
<cjming@deploy1003> |
helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply |
[production] |
| 19:56 |
<cjming@deploy1003> |
helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply |
[production] |
| 19:56 |
<cdobbins@cumin1004> |
END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie |
[production] |
| 19:50 |
<brett@cumin1004> |
END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on A:cp-upload_magru and not P{cp7011.magru.wmnet} and A:cp - 7.1.1-2~bpo13+wmf3 () |
[production] |
| 19:46 |
<vriley@cumin1004> |
START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie |
[production] |
| 19:36 |
<ryankemper> |
[Cirrus] All cirrus pools are serving again. Actively monitoring while the system returns to equilibrium, but all initial indications are that things are as they should be |
[production] |
| 19:34 |
<ebernhardson@deploy1003> |
helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply |
[production] |
| 19:34 |
<ebernhardson@deploy1003> |
helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply |
[production] |
| 19:33 |
<ryankemper@cumin2003> |
END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) pool search-omega in codfw: maintenance |
[production] |
| 19:31 |
<btullis@cumin1004> |
START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet |
[production] |
| 19:31 |
<btullis@cumin1004> |
END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet |
[production] |
| 19:31 |
<btullis@cumin1004> |
START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet |
[production] |
| 19:29 |
<ebernhardson@deploy1003> |
helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply |
[production] |
| 19:29 |
<ebernhardson@deploy1003> |
helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply |
[production] |
| 19:28 |
<ryankemper@cumin2003> |
START - Cookbook sre.discovery.service-route pool search-omega in codfw: maintenance |
[production] |
| 19:27 |
<cdanis@cumin1004> |
conftool action : set/pooled=true; selector: name=codfw,dnsdisc=k8s-ingress-aux-ro |
[production] |
| 19:26 |
<ryankemper> |
[Cirrus] nevermind, that's just the cookbook assuming the DNS record should exist, which it doesn't because chi/psi/omega all share `search.svc.$DC.wmnet` |
[production] |
| 19:25 |
<btullis@cumin1004> |
END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1012.eqiad.wmnet |
[production] |
| 19:24 |
<ryankemper> |
[Cirrus] `dns.resolver.NoAnswer: The DNS response does not contain an answer to the question: search-psi.svc.eqiad.wmnet` checking briefly if this is real failure or just some TTL wonkiness |
[production] |
| 19:23 |
<ryankemper@cumin2003> |
END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) pool search-psi in codfw: maintenance |
[production] |
| 19:20 |
<ebernhardson@deploy1003> |
helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply |
[production] |
| 19:20 |
<ebernhardson@deploy1003> |
helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply |
[production] |
| 19:18 |
<dzahn@cumin2003> |
END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie |
[production] |
| 19:18 |
<ryankemper@cumin2003> |
START - Cookbook sre.discovery.service-route pool search-psi in codfw: maintenance |
[production] |
| 19:17 |
<ryankemper> |
[Cirrus] codfw chi (big cluster) repooled; metrics are already improving, I see poolcounter rejections dropping significantly |
[production] |
| 19:17 |
<ryankemper@cumin2003> |
END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool search in codfw: maintenance |
[production] |
| 19:17 |
<cdanis@cumin1004> |
conftool action : set/ttl=300; selector: name=codfw |
[production] |
| 19:13 |
<cdobbins@cumin1004> |
START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie |
[production] |
| 19:12 |
<ryankemper@cumin2003> |
START - Cookbook sre.discovery.service-route pool search in codfw: maintenance |
[production] |
| 19:11 |
<ryankemper> |
[Cirrus] Repooling codfw, chi first followed by the small clusters |
[production] |
| 19:11 |
<cdanis@cumin1004> |
conftool action : set/pooled=true; selector: name=codfw,dnsdisc=(kartotherian|tegola-vector-tiles) |
[production] |
| 19:07 |
<cdobbins@cumin1004> |
END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie |
[production] |
| 19:02 |
<jhuneidi@deploy1003> |
Finished scap sync-world: Backport for [[gerrit:1344753|REST: restore PageContentHelper::checkAccess (fix live breakage)]] (duration: 10m 15s) |
[production] |
| 18:57 |
<jhuneidi@deploy1003> |
daniel, jhuneidi: Continuing with deployment |
[production] |
| 18:56 |
<jhuneidi@deploy1003> |
daniel, jhuneidi: Backport for [[gerrit:1344753|REST: restore PageContentHelper::checkAccess (fix live breakage)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. |
[production] |
| 18:55 |
<btullis@cumin1004> |
START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet |
[production] |
| 18:55 |
<btullis@cumin1004> |
END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet |
[production] |
| 18:55 |
<btullis@cumin1004> |
START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet |
[production] |
| 18:52 |
<jhuneidi@deploy1003> |
Started scap sync-world: Backport for [[gerrit:1344753|REST: restore PageContentHelper::checkAccess (fix live breakage)]] |
[production] |