NetBird Part 3 - Access Policies, or RBAC for Your Network
NetBird Part 3 - Access Policies
The first two posts got you a self-hosted NetBird mesh: machines find
each other peer-to-peer, coordinated by a control plane you run, and
join with setup keys. But right now that mesh is flat. Every peer can
reach every other peer on every port. Your laptop, your database
server, the little monitoring box, the machine a contractor enrolled
last week, all one happy family with full reach to each other.
That is not a network you want for real infrastructure, and fixing it
is the entire reason to run a control plane instead of hand-meshing
WireGuard. This post is about access policies: turning the flat mesh
into a network where the database is reachable only by the things that
should reach it, and a compromised peer reaches almost nothing. This is
RBAC, applied to network connectivity.
Deny by Default, With One Catch
The good news first: NetBird is deny-by-default at its core. Without any
policy, no peer can talk to any other peer. That is the correct
foundation, the same deny-by-default floor that makes Kubernetes RBAC
sane: nothing is allowed until you say so, and every bit of
connectivity is something you granted on purpose.
The catch is that to make the initial setup painless, NetBird ships a
Default policy that allows every peer to reach every other peer on all
protocols and ports. So out of the box you get the flat mesh, not
because deny-by-default failed, but because there is one broad allow
rule sitting on top of it. The single most important action when you
move from "playing with it" to "depending on it" is to delete that
Default policy. The moment it is gone, you are back to true
deny-by-default, nothing can talk to anything, and you build
connectivity back up deliberately, one policy at a time.
Policies in NetBird are allow-only. There are no deny rules and no
priority ordering to reason about, because you cannot write a deny.
Everything is denied unless some allow policy permits it. This is a
genuinely nice property: there is no rule-order puzzle, no "which rule
wins," just a set of explicit permissions over a default of nothing.
Groups Are the Unit, Not Peers
You do not write policies about individual machines. You write them
about groups, and getting the group model right is the whole game.
A group is just a label you attach to peers, and a policy says one group
may reach another group, on specific protocols and ports. "The app
servers may reach the database group on TCP 5432." "The admin group may
reach everything on everything." The peers are interchangeable; the
groups carry the meaning. This is what keeps the system manageable as it
grows: you add a new app server to the app group and it inherits exactly
the right access with no new policy.
NetBird's own guidance, which is worth following, is to keep two kinds of
group mentally separate even though the system treats them identically.
User groups are the people, used as the source of a policy: who is
allowed to initiate access. Peer groups are the infrastructure, used as
the destination: what is being accessed. The clean pattern is user
groups as source, infrastructure peer groups as destination. Users
access servers; servers do not access users. Keeping that direction
straight is most of what keeps an access model readable a year later.
There is a built-in All group that contains every peer automatically. It
is convenient and it is a trap: a policy with All as source or
destination is the flat mesh creeping back in. Use it sparingly and
deliberately, never as the lazy default.
Setup Keys Assign Groups Automatically
This is where part two's setup keys connect to the policy model, and it
is the piece that makes the whole thing scale.
A setup key can auto-assign groups to any peer that enrolls with it. You
create a key for, say, your database tier, configure it to auto-assign
the database group, and every machine you enroll with that key lands in
the database group automatically, inheriting exactly the policies that
group has, with no manual step. Deploy a new database replica, enroll it
with the database key, and it is governed correctly from the first
second it appears on the network.
This is the mechanism that makes "governed by default" real rather than
aspirational. The alternative, enrolling machines and then remembering
to put each one in the right group by hand, is the kind of manual step
that gets skipped under pressure, and a skipped step here means a machine
sitting in no group, or the wrong one, with the wrong reach. Auto-group
on the setup key removes the human from that loop. Scope one key per
tier or environment, and enrollment and authorization become the same
action.
Two Sharp Edges Worth Knowing
NetBird's policy model has two behaviors that surprise people, and both
can quietly leave you more open than you think.
Peers in the same group can always reach each other. Access policies
govern traffic between different groups, not within a group. If you put
ten machines in one group, those ten can ping, scan, and SSH each other
directly regardless of your policies, because intra-group communication
is not something policies restrict. The practical consequence: group
size is a blast-radius decision. Ten machines in one group means a
compromise of any one of them reaches the other nine freely. If two
machines should not reach each other, they must be in different groups,
full stop. Do not put things in the same group for convenience that you
would not want talking to each other.
Network routes can bypass access control. When you route traffic to a
whole subnet through a routing peer, a route configured without ACL
groups grants unrestricted access to that entire destination CIDR for
every peer that can use the route. From a zero-trust posture that is the
wrong default, and it is easy to hit by accident because the route
"just works" without the ACL groups, so you never notice the door is
wide open. Whenever you set up subnet routing, configure the ACL groups
explicitly, or you have punched a hole that your careful peer-to-peer
policies do not cover.
Posture Checks: Access That Depends on the Device
The step beyond "who may reach what" is "under what condition," and this
is where NetBird earns the zero-trust label. A policy can carry posture
checks: requirements the connecting peer's device must meet before the
policy grants access. The checks available include a minimum NetBird
client version, the operating system and version, and the peer's
geographic location.
The use is straightforward and powerful. Require a current client
version to reach sensitive systems, and a machine running an old,
possibly vulnerable agent is refused rather than trusted. Restrict
access to a tier by geography, and a connection from somewhere it should
never originate is blocked even if the credentials and the group are
right. Access stops being a static yes-or-no tied only to identity and
becomes conditional on the state of the device asking, which is the
actual meaning of zero trust: never trust, always verify, and verify the
context, not just the name.
The Point
A flat mesh where everything reaches everything is easy to stand up and
wrong to depend on. NetBird's access policies turn it into a governed
network, and the model is clean: deny by default once you delete the
starter policy, allow-only rules so there is no ordering puzzle, groups
as the unit of access with users as source and infrastructure as
destination, and setup keys that auto-assign groups so enrollment and
authorization are one action. Posture checks add the conditional layer
that makes it genuinely zero trust.
Delete the Default policy. Design your groups as blast-radius boundaries,
remembering that same-group peers always reach each other. Write
allow-only policies from user groups to infrastructure groups. Lock down
subnet routes with ACL groups so they do not bypass everything else. And
use posture checks on anything sensitive. Do that and the mesh stops
being a flat network with a VPN's worth of implicit trust and becomes
what the control plane was always for: least privilege, enforced at the
network layer, where a compromise reaches only what you decided it could.
The series so far: the concept, the self-hosted control plane, and now
access policies. The mesh is built, it is yours, and it is governed. The
next post gets into routing, reaching whole subnets and the resources
that cannot run the agent themselves, where the ACL-group caveat above
becomes a section of its own.