201-250 of 10000 results (36ms)
2026-07-20 ยง
15:33 <sukhe@cumin1003> cookbooks.sre.cdn.roll-reboot finished rebooting cp2050.codfw.wmnet [production]
15:28 <bking@cumin2003> START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - T431826 [production]
15:24 <cwilliams@cumin1003> dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94906 and previous config saved to /var/cache/conftool/dbconfig/20260720-152418-cwilliams.json [production]
15:14 <bking@cumin2003> END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1023 [production]
15:14 <bking@cumin2003> START - Cookbook sre.hosts.move-vlan for host wdqs1023 [production]
15:14 <bking@cumin2003> START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm [production]
15:14 <cwilliams@cumin1003> dbctl commit (dc=all): 'Repooling after maintenance db2162 (T431660)', diff saved to https://phabricator.wikimedia.org/P94905 and previous config saved to /var/cache/conftool/dbconfig/20260720-151407-cwilliams.json [production]
15:13 <urbanecm@deploy2003> Finished scap sync-world: Backport for [[gerrit:1312471|postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423|Modify user groups rights in English Wikiquote (T432557)]] (duration: 41m 16s) [production]
15:08 <kevinbazira@deploy2003> helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' . [production]
15:07 <cwilliams@cumin1003> dbctl commit (dc=all): 'Depooling db2162 (T431660)', diff saved to https://phabricator.wikimedia.org/P94902 and previous config saved to /var/cache/conftool/dbconfig/20260720-150729-cwilliams.json [production]
15:07 <cwilliams@cumin1003> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2162.codfw.wmnet with reason: Maintenance [production]
15:05 <bking@cumin2003> END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) (T430880, restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards [production]
15:04 <wm-bot2> Deployment completed: https://github.com/cluebotng/component-configs/actions/runs/29753268267 (https://github.com/cluebotng/component-configs/commits/3cd242ed54e9146f6a3dfd9305f97d4580c2fc06) [tools.cluebotng]
15:00 <urbanecm@deploy2003> vadymts1, migr, urbanecm: Continuing with deployment [production]
14:59 <bking@cumin2003> END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply security updates - bking@cumin2003 - T431826 [production]
14:58 <bking@deploy2003> Finished deploy [wdqs/wdqs@e8fb00c]: T430880 (duration: 00m 07s) [production]
14:58 <bking@deploy2003> Started deploy [wdqs/wdqs@e8fb00c]: T430880 [production]
14:58 <bking@deploy2003> Finished deploy [wdqs/wdqs@e8fb00c]: T430880 (duration: 00m 13s) [production]
14:58 <bking@deploy2003> Started deploy [wdqs/wdqs@e8fb00c]: T430880 [production]
14:57 <bking@cumin2003> START - Cookbook sre.wdqs.data-transfer (T430880, restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards [production]
14:57 <sukhe@cumin1003> cookbooks.sre.cdn.roll-reboot finished rebooting cp2047.codfw.wmnet [production]
14:55 <sukhe@cumin1003> cookbooks.sre.cdn.roll-reboot finished rebooting cp2048.codfw.wmnet [production]
14:51 <bking@cumin2003> END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2019.codfw.wmnet with OS bookworm [production]
14:49 <sukhe@cumin1003> END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica and not P{lvs7003*} and A:liberica [production]
14:47 <urbanecm@deploy2003> vadymts1, migr, urbanecm: Backport for [[gerrit:1312471|postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423|Modify user groups rights in English Wikiquote (T432557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there. [production]
14:44 <sukhe@cumin1003> END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified] [production]
14:44 <sukhe@cumin1003> START - Cookbook sre.dns.admin DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified] [production]
14:41 <sukhe@cumin1003> END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart P{lvs7003.magru.wmnet} and A:liberica [production]
14:41 <sukhe@cumin1003> END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P{lvs7003.magru.wmnet} and A:liberica [production]
14:41 <sukhe@cumin1003> START - Cookbook sre.loadbalancer.admin pooling P{lvs7003.magru.wmnet} and A:liberica [production]
14:41 <sukhe@cumin1003> END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P{lvs7003.magru.wmnet} and A:liberica [production]
14:39 <sukhe@cumin1003> START - Cookbook sre.loadbalancer.admin depooling P{lvs7003.magru.wmnet} and A:liberica [production]
14:39 <sukhe@cumin1003> START - Cookbook sre.loadbalancer.upgrade restart P{lvs7003.magru.wmnet} and A:liberica [production]
14:33 <sukhe@cumin1003> END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: no reason specified, no task ID specified] [production]
14:33 <sukhe@cumin1003> START - Cookbook sre.dns.admin DNS admin: depool magru [reason: no reason specified, no task ID specified] [production]
14:31 <urbanecm@deploy2003> Started scap sync-world: Backport for [[gerrit:1312471|postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423|Modify user groups rights in English Wikiquote (T432557)]] [production]
14:24 <jgiannelos@deploy2003> helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply [production]
14:24 <jgiannelos@deploy2003> helmfile [codfw] START helmfile.d/services/mw-parsoid: apply [production]
14:24 <jgiannelos@deploy2003> helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply [production]
14:24 <jgiannelos@deploy2003> helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply [production]
14:23 <bking@cumin2003> END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage [production]
14:22 <bking@cumin2003> START - Cookbook sre.wdqs.data-transfer (T430880, restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards [production]
14:19 <bking@cumin2003> START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage [production]
14:19 <bking@cumin2003> END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2027.codfw.wmnet with OS bookworm [production]
14:16 <sukhe@cumin1003> cookbooks.sre.cdn.roll-reboot finished rebooting cp2046.codfw.wmnet [production]
14:16 <sukhe@cumin1003> cookbooks.sre.cdn.roll-reboot finished rebooting cp2045.codfw.wmnet [production]
14:11 <wm-bot2> Deployment completed: https://github.com/cluebotng/component-configs/actions/runs/29749326972 (https://github.com/cluebotng/component-configs/commits/83d05d6d10640a88d80bdbc041955ca27603a0fe) [tools.cluebotng]
14:08 <jgiannelos@deploy2003> helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply [production]
14:08 <jgiannelos@deploy2003> helmfile [codfw] START helmfile.d/services/mw-parsoid: apply [production]
14:08 <jgiannelos@deploy2003> helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply [production]