| 1 | # 2024-03-07 |
| 2 | |
| 3 | Attendees: hexa, vcunat, zimbatm (Jonas), Linus, Julien, Raito/Ryan, Jade (most |
| 4 | of the time) |
| 5 | |
| 6 | ## [hexa] arm64 hetzner machine config |
| 7 | |
| 8 | - Dump it into a new directory in the infra repo, allow infra-build to deploy |
| 9 | - vcunat: There's an issue containing the bits of the configuration |
| 10 | - vcunat: I assumed we wanted to migrate it directly to a new deployment |
| 11 | system |
| 12 | - hexa: delroth wanted to script out iPXE but this has not panned out yet, we |
| 13 | discovered we had DHCP available, which is promising |
| 14 | |
| 15 | ## [zimbatm] Round table |
| 16 | |
| 17 | What is on everyone's mind? What are your plans? |
| 18 | |
| 19 | - Linus: |
| 20 | - Happy to help out with stuff, pairing on with anything |
| 21 | - zimbatm: Do you think we should do a better presentation? |
| 22 | - linus: I think that'd be good |
| 23 | - hexa: |
| 24 | - Looking at iPXE, hold us back the most right now |
| 25 | - will coordinate with delroth, if he has already anything |
| 26 | - Open to discuss the Ceph scenario |
| 27 | - A lot of discussions ongoing with the self-hosted binary cache, that's |
| 28 | good |
| 29 | - We are running into questions that cannot be answered by anyone |
| 30 | - What should be the availability? |
| 31 | - What should be the durability? |
| 32 | - Discussion running in circles right now |
| 33 | - Form a tightr discussion group |
| 34 | - So that you can identify the main points |
| 35 | - And address them |
| 36 | - And not run into circles |
| 37 | - vcunat: |
| 38 | - Continuously busy with staging iterations |
| 39 | - Unblocking difficult to access machines, e.g. aarch64 machine |
| 40 | - There's actually more of my machines in the infra and that also requires |
| 41 | update |
| 42 | - Small benchmarking machine that makes sense: |
| 43 | - t2a |
| 44 | - The point is to have consistent benchmarking data |
| 45 | - Linus: we definitely don't have cloud VMs for benchmarking, we probably |
| 46 | want dedicated hardware |
| 47 | - zimbatm: could you potentially create a ticket to make an inventory of your |
| 48 | machines? |
| 49 | - vcunat: there's two machines: t2a and t4b only really |
| 50 | - Julien: |
| 51 | - _Short-term_: I would like to onboard more folks on non-critical |
| 52 | infrastructure |
| 53 | - I would like to give them tasks to do end to end |
| 54 | - Difficult to do with the current list of tasks atm |
| 55 | - The wiki is also something I also want to get out ASAP |
| 56 | - The technical issues are basically non-existent, just a little bit more |
| 57 | work to do |
| 58 | - Then announcements, onboard people to do editorial work, and that's it |
| 59 | - We are near ready to launch |
| 60 | - zimbatm: Bitwarden |
| 61 | - Julien: we need to move the data from old to new and inform the change to |
| 62 | the users |
| 63 | - zimbatm: OK, we need to organize that migration |
| 64 | - Julien: we can discuss this async |
| 65 | - Interested also in cache self-hosting discussions |
| 66 | - We have momentum and it'd be nice to have some sort of stance from infra |
| 67 | people |
| 68 | - Addressing the recent unrest regarding the public stance of infra on self |
| 69 | hosting |
| 70 | - zimbatm: we should/could do a proof of concept so we can get a feeling |
| 71 | about how easy is it to operate |
| 72 | - Ryan: |
| 73 | - Recommend https://github.com/zhaofengli/colmena/pull/198 |
| 74 | |
| 75 | Things to pick up for infra: |
| 76 | https://github.com/NixOS/infra/issues?q=is%3Aissue+is%3Aopen+sort%3Aupdated-desc+label%3Anew-service |
| 77 | |
| 78 | ## [hexa] darwin access |
| 79 | |
| 80 | - hexa: We have an inventory problem |
| 81 | - What machines exist? What machines should we be able to access? |
| 82 | - Important so we can delegate access and unblock work |
| 83 | - |
| 84 | |
| 85 | - Braindump |
| 86 | - Apple M1 at Hetzner (hydra) |
| 87 | - Apple M1 in Grahams basement (???) |
| 88 | - Apple M1 at Macstadium (ofborg) |
| 89 | - Apple x86_64 at Macstadium (ofborg) |
| 90 | |
| 91 | ## [hexa] ofBorg access |
| 92 | |
| 93 | - hexa: we have some folks who want to work on OfBorg but cannot do because they |
| 94 | are not empowered on to do so |
| 95 | - it is also go via buildkite management mechanism from Graham |
| 96 | |
| 97 | ## [raito] aarch64.nixos.community management |
| 98 | |
| 99 | - https://github.com/NixOS/aarch64-build-box/ |
| 100 | - managed by community or infra? |
| 101 | - zimbatm: it used to be in the nix-community infra, but because the |
| 102 | nix-community does not have access to the Packet account |
| 103 | - hexa: in the past, the worst we had is to debug the kernel issues, which is |
| 104 | difficult w/o packet access |
| 105 | - utilized by ofBorg, too, not a problem because we don't need to trust its |
| 106 | build results |
| 107 | - zimbatm: will talk with zowoq, who manages the nix-community day-to-day |
| 108 | operation |
| 109 | |
| 110 | ## Changelog |
| 111 | |
| 112 | - Cancelled the contract for `eris.nixos.org` (ends after 2024-02-28) |
| 113 | - All services have been migrated to pluto.nixos.org |
| 114 | - Set up backups for Prometheus, Grafana, VictoriaMetrics |
| 115 | - The primary hostnames for Prometheus and Grafana have changed |
| 116 | - https://prometheus.nixos.org |
| 117 | - https://grafana.nixos.org |
| 118 | - Redirects for the old hostname/path are in place |
| 119 | - Hydra changes |
| 120 | - Increase pipe size to improve queue-runner performance |
| 121 | - Increased retention interval of Prometheus to two years so we have more |
| 122 | history to evaluate these changes |
| 123 | - Builders have received the fix for |
| 124 | https://github.com/NixOS/nix/security/advisories/GHSA-2ffj-w4mj-pg37 |
| 125 | - GitHub App for wiki.nixos.org so users can log in. |