Each protection entry is has a lower and upper bound. The entry can cover something as small as a single byte, or as big as the whole address space or anything inbetween.
It is just a normal bounds compare to see if each entry covers the access, and the PLB has as long as the top-level cache access takes to do the checks.
So the PLB misses far far less often than a conventional TLB.
TLBs index a tag with the higher bits and compare one retrieved value for equality only. A PLB with arbitrary resolution would need to do 2 subtractions for the 2 compares and do that for all active entries in parallel. That is, for every simple load/store in a 16 entry PLB you'd need to do 32 subtractions! Unless you come up with a novel scheme to handle this your chip will get hot. Maybe you can reconstruct the indexing TLB scheme on the fly in hardware or something like that, or have an special purpose embedded Cortex-M3 that recreates an efficient lookup structure for the MILL :-). And no, a 64 bit subtraction is not one clock cycle when you run at 1ghz.
It is just a normal bounds compare to see if each entry covers the access, and the PLB has as long as the top-level cache access takes to do the checks.
So the PLB misses far far less often than a conventional TLB.