Sampling Evolution¶
Are the flow caches effective to support IPFIX implementations today? What happens if we stop using them? Learn about the new IPFIX implementation in Juniper PTX and ACX7000 routers.
Introduction¶
NetFlow was initially designed as a flow monitoring technique where the network element monitors IP flows, aggregates statistics over multiple packets of the same flow, and periodically exports flow statistics to the collector. Flow statistics typically include information that identifies a flow (a combination of IP header fields), and metadata, such as the incoming and outgoing interfaces of the packet. Modern versions of the NetFlow protocol are standardized in IETF, RFC 5101 and RFC 7011, and it is now called IPFIX. These protocols are extensible and allow reporting of arbitrary information from the network element. IPFIX applications extend beyond traditional routing applications: IPFIX is used for Carrier Grade NAT session logging, and arbitrary MIB statistics export.
However, the main IPFIX application is the export of flows, and IPFIX is strongly associated with flow aggregation on the router itself.
This article aims to demonstrate that there is little to no benefit in using flow aggregation on the router anymore for many practical core and peering applications. More than that, there are applications where flow aggregation is not desirable, because of the extra latency aggregation it introduces.
Note: This article has been written with PTX10000 in mind, but the same implementation is used in the ACX7000 Series too. The MX Series is still using a flow cache approach.
Flow Export Implementation¶
A typical IPFIX flow export implementation is comprised of two processes, export and sampling, plus a flow table where millions of flows reside temporarily.

When a packet is sampled, the sampling process extracts packet header fields that form the key and consults the flow table. In case of a miss, a new entry is added, and the flow key and other fields of the new entry are initialized. If an existing flow table entry is hit, the packet and byte count fields of an existing flow are incremented.
The export process periodically walks the table and exports flows that had no activity for the specified duration (inactive flow timeout), or active flows that have not been exported for a long time (active flow timeout).
Sampling and export processes may be implemented in the ASIC, or in software. Flow table may also reside in the ASIC or in software. Hybrid models are possible too. Regardless of the implementation, how effective is the flow aggregation?
The answer depends on the sampling rate, and timeouts configuration. Figure 2 shows statistics collected from the production network serving Internet traffic. With 1:4096 sampling rate, 60s active flow timeout and 15s inactive flow timeout 90.7661% of flow records contain only one packet.

To measure the effectiveness of the flow table we need to measure the proportion of packets that hit known flow table entries. The formula to compute this number is rather simple and it is shown on the plot. Only 16.56% of packets can leverage the flow aggregation capability of the device in this case.
Both total packet statistics and total flow statistics is reported by JUNOS today and can be collected by running the command below.
We conducted analysis on several deployments, and the results are presented in the Table 1.
| ID | Rate | Inactive Timeout | Active Timeout | Flow Table Hit Ratio |
|---|---|---|---|---|
| 1 | 100 | 15 | 600 | 80.01% |
| 2 | 500 | 60 | 60 | 63.41% |
| 3 | 1000 | 15 | 60 | 46.13% |
| 4 | 1000 | 15 | 60 | 55.86% |
| 5 | 1000 | 60 | 60 | 35.38% |
| 6 | 1024 | 60 | 1800 | 16.40% |
| 7 | 2000 | 60 | 1800 | 1.74% |
| 8 | 4000 | 60 | 1800 | 29.43% |
| 9 | 4096 | 60 | 1800 | 17.66% |
| 10 | 8192 | 60 | 120 | 13.55% |
| 11 | 10000 | 60 | 1800 | 16.39% |
| 12 | 32767 | 60 | 1800 | 3.54% |
Table 1: Flow Table Hit Ratio
As the data suggests, the percentage of flows with a single packet increases with the sampling rate. And there is a theoretical explanation for this.
Figure 3 shows the theoretical analysis of single packet flow percentage as function of the sampling rate (from 16 to 8192) and different flow lengths (from 16 to 4096), see the Appendix section for details how these plots are produced.

Studies in Annex [1] and [2] show that more than 99% of the flows have less than 200 packets, hence most of the time only single packet from the flow is sampled with reasonable sampling rates.
If only single packet is sampled, then flow aggregation table maintenance offers no benefit, and the implementation can be simplified.
In a simplified implementation, flow table with millions of entries is replaced with a small circular buffer with hundreds of entries, Figure 4.
Sampling process adds new entries to the buffer, and export process gathers them, creates flow reports and sends out to the collector.

Conclusion¶
Simplified IPFIX sampling implementation reduces memory footprint and increases the performance: no flow table management is needed. There is another added benefit: flow reporting latency reduces to less than a second, compared to 15 and more seconds in IPFIX case (typical active flow and inactive flow timeouts). If no flow table is managed, the next step in the evolution process is to report the packet content instead of just packet fields. This is a topic for the future article. Keep in mind that sampling rate must be relatively high for the approach to be feasible, but these sampling rates are typical in most of the peering and core deployments. Lower sampling rates, down to 1:1, require very different implementation, to be described in the future article.
Application in Juniper Routers¶
Starting from 21.3 release, Juniper 100GE PTX routers based on Express 2 chipset (Paradise) use new simplified implementation by default, with an option to fall back to the flow cache mode when nexthop-learning knob is configured, see documentation. 400GE PTX routers only use new simplified IPFIX implementation, and demonstrate outstanding sampling performance among routers in its class, up to 150 thousands of sampled packets per second per line card or a fixed form factor system.
Appendix: Theoretical Analysis R Script¶
The R script below is used to produce the Figure 3.
References¶
- Qian and B. E. Carpenter, "A flow-based performance analysis of TCP and TCP applications," 2012 18th IEEE International Conference on Networks (ICON), 2012, pp. 41-45, doi: 10.1109/ICON.2012.6506531.
- Jurkiewicz, Piotr, Grzegorz Rzym, and Piotr Borylo. "Flow length and size distributions in campus Internet traffic." Computer Communications 167 (2021): 15-30
- https://datatracker.ietf.org/doc/html/rfc5101
- https://datatracker.ietf.org/doc/html/rfc7011
Glossary¶
- ASIC: Application Specific Integrated Circuit
- CGNAT: Carrier Grade Network Address Translation
- IETF: Internet Engineering Task Force
- RFC: Requests for Comments
- IPFIX: IP Flow Information Export
Acknowledgements¶
Many thanks to Ranjith S V V Kumar and Alex Baban for reviewing this article.