1-50 of 10000 results (104ms)
2026-07-25 §
00:11 <ryankemper> [WDQS] T430880 Repooled wdqs1014.eqiad.wmnet and wdqs2008.codfw.wmnet after Bookworm reimage, transfer, and postflight; wdqs2008 is serving, while wdqs1014 will remain outside of service until a pybal restart next monday [production]
2026-07-24 §
23:54 <ryankemper@cumin2003> conftool action : set/pooled=yes; selector: name=wdqs1014.eqiad.wmnet [production]
23:54 <ryankemper@cumin2003> conftool action : set/pooled=yes; selector: name=wdqs2008.codfw.wmnet [production]
23:43 <jhathaway@cumin1003> END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2010.codfw.wmnet with OS trixie [production]
23:08 <jhathaway@cumin1003> END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage [production]
23:03 <jhathaway@cumin1003> START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage [production]
22:33 <jhathaway@cumin1003> START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie [production]
22:13 <ssastry@deploy1003> helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply [production]
22:13 <ssastry@deploy1003> helmfile [codfw] START helmfile.d/services/mw-parsoid: apply [production]
22:13 <ssastry@deploy1003> helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply [production]
22:13 <ssastry@deploy1003> helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply [production]
22:00 <jhathaway@cumin1003> END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie [production]
21:53 <jhathaway@cumin1003> START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie [production]
21:53 <jhathaway@cumin1003> END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie [production]
21:51 <jhathaway@cumin1003> START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie [production]
21:47 <jhathaway@cumin1003> END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie [production]
21:43 <jhathaway@cumin1003> START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie [production]
21:39 <jhathaway@cumin1003> END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie [production]
21:38 <jhathaway@cumin1003> START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie [production]
17:21 <jiji@cumin1003> END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1135.eqiad.wmnet [production]
17:21 <jiji@cumin1003> END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1135.eqiad.wmnet [production]
17:21 <jiji@cumin1003> START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1135.eqiad.wmnet [production]
16:34 <bd808@deploy1003> helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply [production]
16:34 <bd808@deploy1003> helmfile [eqiad] START helmfile.d/services/developer-portal: apply [production]
16:34 <bd808@deploy1003> helmfile [codfw] DONE helmfile.d/services/developer-portal: apply [production]
16:34 <bd808@deploy1003> helmfile [codfw] START helmfile.d/services/developer-portal: apply [production]
16:33 <bd808@deploy1003> helmfile [staging] DONE helmfile.d/services/developer-portal: apply [production]
16:33 <bd808@deploy1003> helmfile [staging] START helmfile.d/services/developer-portal: apply [production]
16:28 <ssastry@deploy1003> helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply [production]
16:28 <ssastry@deploy1003> helmfile [codfw] START helmfile.d/services/mw-parsoid: apply [production]
16:28 <ssastry@deploy1003> helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply [production]
16:28 <ssastry@deploy1003> helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply [production]
16:11 <ssastry@deploy1003> helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply [production]
16:11 <ssastry@deploy1003> helmfile [codfw] START helmfile.d/services/mw-parsoid: apply [production]
16:11 <ssastry@deploy1003> helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply [production]
16:11 <ssastry@deploy1003> helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply [production]
16:03 <bking@cumin2003> END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) (T430880, restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2008.codfw.wmnet, repooling source-only afterwards [production]
16:03 <bking@cumin2003> END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) (T430880, restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1014.eqiad.wmnet, repooling source-only afterwards [production]
15:56 <jiji@cumin1003> END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1135.eqiad.wmnet with OS trixie [production]
15:54 <elukey@cumin1003> END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 40 hosts [production]
15:52 <elukey@cumin1003> START - Cookbook sre.hosts.bmc-user-mgmt for 40 hosts [production]
15:37 <topranks> upgrade SR-Linux OS on lswtest-d8-eqiad [production]
15:36 <jiji@cumin1003> END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1135.eqiad.wmnet with reason: host reimage [production]
15:33 <cmooney@cumin1003> DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 6 hosts with reason: upgrade lswtest-d8-eqiad [production]
15:32 <jiji@cumin1003> START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1135.eqiad.wmnet with reason: host reimage [production]
15:30 <jiji@cumin1003> END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc-gp2006.codfw.wmnet with OS bookworm [production]
15:15 <jiji@cumin1003> END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1135 [production]
15:15 <jiji@cumin1003> END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1135 [production]
15:13 <jiji@cumin1003> END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc-gp2006.codfw.wmnet with reason: host reimage [production]
15:08 <jiji@cumin1003> START - Cookbook sre.hosts.downtime for 2:00:00 on mc-gp2006.codfw.wmnet with reason: host reimage [production]