[–] 3 points 1 week ago

LineageOS, mostly due to the fact that it's pretty close to the original vision of AOSP; and my devices do not support Graphene or PostmarketOS, so there are almost no viable alternatives.

  • source
  • [–] [S] 1 point 3 weeks ago

    That does seem to be a good solution, thank you for the recommendation!

    Running an AS and obtaining a /48 through a sponsoring LIR seems to cost about $150/year in my country, but hardly anyone (except hosting providers and large companies) does this, since self-hosting, especially more complicated stuff, isn't really that popular here; people are mostly uninformed that it even exists.

    But route64 is apparently completely free (donations are welcome); they provide a /56 PA-like IPv6 block carved out of their PI, and they also handle BGP stuff, and I don't even have to pay for any of this - overall, a great choice, considering my constraints. Anycast to multiple tunnels costs money (maybe that's what you meant with 2 EUR/month?) but that's pretty much lunch money so it would be OK with me.

  • source
  • parent
  • context
  • [–] [S] 4 points 3 weeks ago

    At this point I pretty much think this may be the exact case - and it may work in reverse: "since ns2 has AAAA listed then it's an IPv6-capable deployment and we are allowed to fall back to IPv4, but since ns1 lacks the AAAA record, then we won't even consider falling back to ns1 if ns2 is down".

    Or maybe another scenario:

    • let's check ns2's AAAA record
    • oh it's down! Let's check ns2's A record
    • it's down too! Let's switch to ns1...
    • ... but at that time the DNS resolution timeout expires, since we took too long resolving the records; and on the second try, no one remembers that ns2 was down, repeating the cycle
  • source
  • parent
  • context
  • [–] [S] 2 points 3 weeks ago

    Yeah, the issue is that people's resolvers are wildly different and I guess some of them do not cater to weird and wacky self-hosting setups (most people just use anycast highly-available DNS, usually provided by a registrar or by some other company).

    I had a hunch that resolvers break because one of my nodes does not have a public IPv6 address: meaning,

    • node2: externally-facing node, I host it from my home, IP 178.44.116.85, acts as NS ns1. It has IPv6 connectivity but no external IPv6;
    • node4: a cloud VPS that I rent. Externally-facing, with IPv4 91.219.150.30 and IPv6 2a06:dd00:1:4::4189, acts as NS ns2.

    If node4 decides to break, then ns2 becomes unavailable - but since resolvers can see that ns1 does not have an IPv4 address, and since everyone knows that IPv6 is better because 6 is greater than 4, then these resolvers would always prefer ns2, even though it's clearly down.

    This could be solved by asking my ISP to set up a static IPv6 prefix for me (they already provide static IPv4), except my ISP does not do that, unfortunately. I could also rent another VPS (so I would have two VPS'es, giving me some resilience from downtimes) but that means extra $$$... anyway, I guess I should probably test the IPv6-preference hypothesis first and then act on the results.

  • source
  • parent
  • context
  • [–] [S] 3 points 3 weeks ago (3 children)

    Basically my internal DNS zonefile is configured like this (it should be easier to read than my explanation):

    coredns_config:
      zoneFiles:
        - filename: int.pootis.network.zone
          domain: int.pootis.network
          contents: |
            $ORIGIN int.pootis.network.
            $TTL 300
    
            @       SOA     ns.int.pootis.network. admin.pootis.network. (
              2026082001
              1200
              300
              1209600
              300
            )
    
            @       NS      ns.int.pootis.network.
            ns      AAAA    {{ k8s_dns_ipv6 }}
            ns      A       {{ k8s_dns_ipv4 }}
    
            vault     CNAME   internaltraefik.internaltraefik.svc.{{ k8s_cluster_domain }}.
            cinny     CNAME   internaltraefik.internaltraefik.svc.{{ k8s_cluster_domain }}.
            grafana   CNAME   internaltraefik.internaltraefik.svc.{{ k8s_cluster_domain }}.
            stalwart  CNAME   internaltraefik.internaltraefik.svc.{{ k8s_cluster_domain }}.
            lemmy     CNAME   internaltraefik.internaltraefik.svc.{{ k8s_cluster_domain }}.
            mail      CNAME   stalwart.stalwart.svc.{{ k8s_cluster_domain }}.
            webmail   CNAME   internaltraefik.internaltraefik.svc.{{ k8s_cluster_domain }}.
    
            ; ACME challenges
            ; int.pootis.network + *.int.pootis.network
            _acme-challenge      CNAME 1ca12d2c-8034-4a09-a940-dc59bd4f38d8.acme-dns.pootis.network.
            ; mail.int.pootis.network
            _acme-challenge.mail CNAME d175b66d-909f-46ba-8572-636ebc235d4c.acme-dns.pootis.network.
    
  • source
  • parent
  • context
  • [–] [S] 1 point 3 weeks ago (1 child)

    Trying to understand your setup: this seems like a split-horizon DNS, where app.example.com is only visible from your internal network (or Tailnet), right?

    If so - I explained my similar setup in this comment; in short, my internal DNS works perfectly because the internal DNS server has a static "service IP" (handled by kube-proxy); if my currently-running pod breaks, k8s will simply spawn a new one and the DNS server IP won't even have to change - meaning, the internal DNS zone continues to be served no matter which node breaks (it'll even survive multiple node failures).

    But the external DNS is different and I can't simply plug in a "fake loadbalancer IP" there... (there are valid workarounds but they either are incredibly finicky or I have to spend an obscene amount of money to implement them)

  • source
  • parent
  • context
  • [–] [S] 2 points 3 weeks ago (4 children)

    I actually already have a local DNS resolver... well, sort of: most of my cluster's internal services, such as Vaultwarden and other stuff, are available only over an in-cluster IKEv2 VPN.

    What I mean is, I have to connect to the VPN first, which pushes a Configuration Payload to my IKEv2 client, and that payload contains a DNS server IP, and that DNS server (exposed by a Kubernetes Service) serves a custom "internal" zone (.int.pootis.network), and proxies everything else either to k8s coredns, or to upstream DNS forwarders.

    Okay, that explanation may have been somewhat complicated (it sounded simpler in my head) but the end result is that I can resolve and reach stuff like lemmy.int.pootis.network (my Lemmy frontend) only while connected to the in-cluster VPN. It is completely hidden from external users (no CT log record, no external DNS record, and only proxyable by internal Traefik instance).

    The issue only occurs with my externally-facing DNS, which is needed to, for example, to reach my Lemmy backend instance (lemmy.pootis.network) and the pictrs deployment. But yeah, internal DNS works perfectly for me and it easily survives a node failure.

  • source
  • parent
  • context
  •  

    I'll just provide my own example: my homelab consists of 6 Kubernetes nodes placed across the country. Some differ by ISP, some are placed in different cities, one is hosted on a cloud provider. Basically it's a very cheap variant of geo-replicating my workloads.

    Two of these nodes are visible from the Internet and have a static IP address; one node also has an IPv6 address. Each node hosts an authoritative DNS server (CoreDNS) for my personal domain pootis.network; and the .network TLD has glue records which point to IPs of these two nodes. This is a classic "self-hosted DNS" scenario.

    Here's an excerpt from my zonefile so you can understand the setup better:

    $ORIGIN pootis.network.
    $TTL 300
    
    @       SOA     ns1.pootis.network. admin.pootis.network. (
      2026082001
      1200
      300
      1209600
      300
    )
    
    ; Nameservers and glue records
    @       NS      ns1.pootis.network.
    @       NS      ns2.pootis.network.
    ns1     A       178.44.116.85
    ns2     A       91.219.150.30
    ns2     AAAA    2a06:dd00:1:4::4189
    

    This 5-record block (NS/A/AAAA) is mirrored into the .network zone by my domain registrar (plus DS for DNSSEC but that's another thing).

    As such, my DNS becomes fully independent - and, in theory, if one of my externally-facing nodes breaks, let's say ns1, then DNS resolvers all over the world (forwarders, recursive, and such) will fall back to ns2, and everything will keep working. Kubernetes will also reorganize the pod placement so all my workloads are available again after a slight downtime.

    That would have been great, if it worked as described, but apparently, after one nameserver in my zone fails, then the resolvers... just give up? Let's say ns1 failed but ns2 is working. The parent zone still points to both nameservers. My external resource records (websites and other stuff) at this point would have already been auto-reconfigured by a custom k8s controller to point to the IP addresses of the node that hosts ns2. Simplifying: the entire world basically sees this after ns1 fails and after TTL caches expire:

    ; all of this has very low TTL, 5 minutes or so
    
    @       NS      ns1.pootis.network. ; from .network 
    @       NS      ns2.pootis.network. ; from .network
    
    ns1     A       178.44.116.85 ; broken. Either from .network glue or from my auth DNS
    ns2     A       91.219.150.30 ; either from .network glue or from my auth DNS
    ns2     AAAA    2a06:dd00:1:4::4189 ; same
    
    ; my-website     A       178.44.116.85 ; does not appear because ns1 is broken- my LB already removed it from the set
    my-website     A       91.219.150.30 ; fronted by a pair of CNAMEs due to loadbalancing but still
    my-website     AAAA    2a06:dd00:1:4::4189 ; same
    

    But even if I query 1.1.1.1 directly for my-website's record, it just doesn't work most of the time because the resolver pins itself to ns1 which is currently failing, or it selects ns1 and does not even care to try ns2.

    To be precise: some resolver implementations DO fall back to ns2 as expected, but most of them just pin themselves to ns1 and then outright refuse to resolve the records in my zone.

    And there's actually no reasonable way out, as far as I can see:

    • moving my DNS infra somewhere else (CloudFlare, for example) is unacceptable since I would like for my homelab to be as independent as practically possible;
    • anycasting, or running a fully-fledged BGP AS is also impossible because that costs a lot of money and I'd like for my homelab to fit into a $10/month budget with room to spare;
    • "live-patching" the NS and glue records in the parent zone (.network), to keep up with the set of my working nodes, is possible, but very unwieldy and somewhat hard to accomplish.

    There's a lot of custom machinery that keeps my workloads running and accessible after a node failure, but all of this becomes completely moot when authoritative DNS is the bottleneck.

    Has anyone been running a similar stack and encountered this problem? I'm aware that the answer is usually "host your DNS at CloudFlare" or "use the registrar's DNS infra" but still...

    [–] [S] 12 points 1 month ago (2 children)

    The enclosure itself is fully DIY (I just yanked out the vape internals and glued the board edges to the vape); there's a hole with an SMA antenna connector, visible on the photo. Plus, there are a few I2C sensors inside the vape, also glued to its inner plastic "walls".

    The hardware is partly-DIY; it's a Meshtastic node on nRF52840 combined with HT-RA62 on a Faketec-like PCB. It runs mainline Meshtastic firmware. I originally planned to make it export its sensor and mesh data over Bluetooth (so my other stuff would ingest the data and put it into my observability pipeline), but decided against it and left this "vape" as a pure Meshtastic node that also announces its temperature, pressure and humidity. It has no screen unlike many other similar nodes.

    Basically it was a fun project, and I even got a good deal on the components (it cost me like $8 total or maybe a little more).

  • source
  • parent
  • context
  • view more: next ›