PCI Express (PCIe) bus: generations, speeds and design
PCIe: the backbone of high-performance systems
PCI Express (PCIe) is the point-to-point serial bus that has become the backbone of modern high-performance electronic systems. It connects critical components - NVIDIA or AMD GPUs, NVMe SSDs, FPGAs (Xilinx/Intel), Intel Xeon SoCs and expansion cards - with throughput ranging from 250 MB/s per lane (PCIe 1.0) to 16 GB/s per lane (PCIe 7.0).
Most high-performance electronics projects involve a PCIe interface. After integrating PCIe links into many custom motherboards over more than 10 years - real-time acquisition systems, embedded AI platforms, edge-computing servers - we have watched this technology jump from 2.5 GT/s (PCIe 1.0) to 64 GT/s (PCIe 6.0), a 25x multiplier in less than 20 years.
TL;DR
- Peripheral Component Interconnect Express (PCIe) is a point-to-point serial bus maintained by the PCI Special Interest Group (PCI-SIG), which specifies the physical, link and transaction layers from PCIe 1.0 through PCIe 7.0.
- Each generation doubles the per-lane throughput: PCIe 3.0 = 8 Gigatransfers per second (GT/s), PCIe 4.0 = 16 GT/s, PCIe 5.0 = 32 GT/s, PCIe 6.0 = 64 GT/s and PCIe 7.0 = 128 GT/s, the last two using Pulse Amplitude Modulation 4 (PAM4) and Flow Control Unit (FLIT) encoding.
- Documented in the official architecture guides from Intel and AMD (Intel PCIe architecture, AMD developer hub), PCIe 5.0 is the de facto standard for datacenter and workstation platforms in 2026.
- On our test bench we have measured a 20 to 30 % gain in vertical eye-diagram opening simply by applying systematic back-drilling to PCIe Gen 4 vias running at 16 GT/s.
PCI Express (Peripheral Component Interconnect Express) is a point-to-point serial bus interface used to connect high-performance internal components such as GPUs, NVMe SSDs and expansion cards. Unlike older parallel buses, PCIe relies on independent lanes that deliver up to 8 GB/s per lane in version 6.0, with low latency and guaranteed backward compatibility.
At AESTECHNO, we design and integrate custom PCIe systems for industrial, medical and AI applications. Our expertise covers high-speed routing, signal-integrity constraints and EMC certification - all critical when a single PCIe x16 link can carry up to 128 GB/s of data.
In this article we share what we have learned in the lab: PCIe architecture and lanes, the evolution of the spec, the design pitfalls to avoid, and our recommendations for picking the right generation for a given application.
Project with a PCIe interface?
Our engineers analyse your requirements and recommend the optimal architecture: PCIe 4.0/5.0/6.0, lane count, routing constraints.
Free 30-minute audit - response within 48h
PCIe bus speed by generation: from 1.0 to 7.0
PCIe bus speed is quoted two ways, and mixing them up is the most common source of confusion. The per-lane rate is the raw signalling speed of a single lane, from 2.5 GT/s on PCIe 1.0 to 128 GT/s on PCIe 7.0. Encoding overhead then turns that signalling rate into usable bandwidth: PCIe 4.0's 16 GT/s works out at roughly 2 GB/s per lane. The link speed is what the device actually gets, that per-lane bandwidth multiplied by the number of lanes in the slot, so a PCIe 4.0 x4 NVMe drive sees about 8 GB/s per direction. Do the same for the next generation and PCIe 5.0's 32 GT/s give about 4 GB/s per lane, so a x16 accelerator reaches about 64 GB/s per direction. Always state a PCIe speed as generation plus lane width, never as a generation alone.
The table below is the lookup that links per-lane throughput, x16 aggregate bandwidth and typical workloads. Per the PCI-SIG, every new generation doubles the per-lane data rate of the previous one while preserving full backward compatibility, so the logic never changes as an x16 slot climbs from roughly 4 GB/s per direction on a Gen 1 link to 256 GB/s per direction on a Gen 7 link. Reading the table serves one precise trade-off: pick the lowest generation that satisfies the need, because every generation step is paid for in signal-integrity constraints, PCB cost and validation effort.
| Version | Raw rate | Per lane | x4 aggregate | x16 aggregate | Typical workload |
|---|---|---|---|---|---|
| PCIe 1.0 | 2.5 GT/s | 250 MB/s | 1 GB/s | 4 GB/s | Legacy, industrial I/O cards |
| PCIe 2.0 | 5 GT/s | 500 MB/s | 2 GB/s | 8 GB/s | Network controllers, instrumentation |
| PCIe 3.0 | 8 GT/s | 1 GB/s | 4 GB/s | 16 GB/s | Mid-range GPU, mainstream SSD |
| PCIe 4.0 | 16 GT/s | 2 GB/s | 8 GB/s | 32 GB/s | High-end GPU, Gen4 NVMe |
| PCIe 5.0 | 32 GT/s | 4 GB/s | 16 GB/s | 64 GB/s | Datacenter, AI/HPC, Gen5 NVMe |
| PCIe 6.0 | 64 GT/s | 8 GB/s | 32 GB/s | 128 GB/s | Hyperscale cloud, AI accelerators |
| PCIe 7.0 | 128 GT/s | 16 GB/s | 64 GB/s | 256 GB/s | Hyperscale AI, 800G networking (spec June 2025) |
Our recommendation: for any new design starting in 2025, target PCIe 4.0 as a minimum. PCIe 5.0 is becoming the standard for bandwidth-hungry workloads (AI training, large-scale storage). PCIe 6.0 is still confined to hyperscale datacenters.
PCIe lab case studies
Our hands-on PCIe experience covers many critical configurations. Here are three representative cases that illustrate the signal-integrity and stack-up trade-offs we routinely face:
- Case 1: PCIe Gen 4, post-layout eye closure caused by an un-back-drilled via stub. At 16 GT/s, a residual 0.5 mm stub creates a resonance inside the useful band and collapses the horizontal margin. Counter to the reflex of swapping the laminate, we first push for systematic back-drilling of the differential-lane transition vias, with a typical 20 to 30 % gain in vertical eye opening.
- Case 2: industrial Intel i5 platform with a heavy PCIe footprint (up to 16 combined Gen 4 / Gen 5 lanes) feeding accelerators and NVMe storage. Stack-up choice is structural here: we favour a hybrid build with high-Tg FR-4 for power and digital layers and a Megtron 6 core for the layers carrying the PCIe lanes. Contrary to the assumption that Gen 5 mandates exotic Tachyon-class material, a clean Megtron 6 design holds up on short Gen 5 links (under 10 cm of cumulative trace).
- Case 3: industrial qualification to IPC-6012 Class 2. On long-life industrial products, PCIe lanes must survive aggressive thermal cycling. We specify IPC-6012 Class 2, with localised Class 3 reinforcement on critical vias in the highest-density zones.
PCIe tooling, standards and materials
Our PCIe methodology rests on a precise set of standards: the PCI-SIG Base Specification v4.0 and v5.0 for the electrical and protocol layers, JEDEC SFF (EDSFF E1.S/E3.S) for the M.2 / U.2 / EDSFF form factors, IPC-6012 Class 2/3 for PCB qualification and IPC-2221/2222 for design rules. As specified by PCI-SIG in v6.0, PAM4 + FLIT encoding reaches 64 GT/s without doubling the Nyquist frequency. On the operating-system side, the Linux kernel PCI documentation (kernel.org) describes PCIe enumeration and Active State Power Management (ASPM). The PCI Express article on Wikipedia offers a thorough generational history, while the IEEE 802.3 family published by IEEE formalises the network pairings often co-integrated with PCIe (e.g. 100 GbE SmartNICs). For industrial connectivity we routinely work with Samtec high-speed connectors and Quectel modules in extended M.2 form factors.
Material selection follows a clear hierarchy keyed to the target generation: high-Tg FR-4 is acceptable up to PCIe Gen 3, Isola I-Speed or 370HR/IS410 for production-grade Gen 3 / Gen 4 builds, Megtron 6 for nominal Gen 4 / Gen 5, and Megtron 7 or Tachyon 100G for long Gen 5 and Gen 6 links. On the validation side we run ANSYS SIwave in post-layout for S-parameter extraction, crosstalk analysis and eye-diagram generation before release to fab. That step has caught more than one via stub or impedance discontinuity that the Altium DRC let through.
Contrary to the belief that PCIe Gen 5 always requires an exotic stack-up, a careful high-Tg FR-4 build can carry Gen 3 with short, well-routed and simulation-validated lanes. In our lab, we have found that the dominant factor is rarely the laminate's permittivity - it is the discipline applied to transition vias, differential-pair symmetry and the quality of return planes along the entire signal path.
What is the PCI Express bus?
The PCI Express bus is a point-to-point serial communication standard used to connect high-performance internal components in an electronic system. Maintained by PCI-SIG, it replaces older parallel buses by offering faster transfers, lower latency and a scalable architecture built around independent lanes. Every peripheral gets its own dedicated link to the controller, unlike the shared buses of earlier generations where bandwidth was divided among all components. That architecture is what lets NVMe SSDs, GPUs and network cards stack up without fighting over the same channel.
The PCIe bus connects internal components in a computer such as:
- Graphics cards
- Network cards
- Expansion cards
- SSDs and other high-performance peripherals
Unlike older parallel buses or simpler serial buses such as I2C and SPI, PCIe adopts a point-to-point serial architecture. The result is faster, more efficient transfers with reduced congestion and latency.
How does the PCI Express bus work?
PCIe operation is a point-to-point topology in which each lane is a full-duplex differential pair able to transmit data in both directions simultaneously. The lane count (x1, x4, x8, x16) determines the total bandwidth available to a given peripheral. Each differential pair carries an encoded signal that embeds its own clock, which eliminates the synchronisation problems of parallel buses. Aggregation is transparent: an x4 link spreads packets across four lanes and reassembles them on arrival. For the board designer, that logical elegance has a physical price: at these rates every trace becomes a transmission line whose impedance, length and layer transitions must be controlled.
The total bandwidth of the bus depends on the lane count in use:
- x16: 16 lanes, mainly used for graphics cards or other demanding accelerators.
- x1: 1 lane, for basic peripherals.
- x4: 4 lanes, ideal for SSDs and many expansion cards.
- x8: 8 lanes, for higher bandwidth needs.
PCIe connector form factors
PCIe connectors are standardised mechanical interfaces (x1, x4, x8, x16), each sized for a specific lane count. The mechanical compatibility allows a smaller card to plug into a larger slot, which gives system builders real installation flexibility. An x16 slot therefore accepts an x1, x4 or x8 card, which runs at the lane count it owns. Beyond classic slots, the standard also comes as M.2 for compact storage, U.2 for server bays and mini-PCIe or M.2 Key E for radio modules: different form factors sharing one protocol.
The most common form factors are:
- PCIe x1: for simple expansion cards.
- x4, x8, x16: for higher performance, especially graphics cards and high-throughput SSDs.
PCIe cards can plug into smaller-sized connectors as well, giving system integrators a great deal of flexibility when laying out their components.
The main advantages of PCIe for your systems
The advantages of PCIe boil down to four technical properties that have made it the dominant standard for high-performance internal interconnects: scalability, high bandwidth, point-to-point architecture and hot-plug. Together they explain its universal adoption in modern systems.
Scalability: depending on your needs, you can choose cards and connectors sized for either modest or extremely demanding bandwidth profiles.
High bandwidth: PCIe delivers transfer rates far above legacy parallel buses.
Point-to-point architecture: this topology cuts latency and improves overall system responsiveness.
Hot-plug: PCIe lets you add or remove peripherals without powering the system down, which simplifies maintenance and upgrade operations.
PCI Express version history: release dates and adoption
The PCI Express version history is the succession of revisions PCI-SIG has released since 2003, each one roughly doubling per-lane bandwidth while staying backward compatible. What matters when you are choosing a generation is less the raw number than the release date and how far adoption has actually travelled: silicon, chipsets and validation tooling lag every specification by several years, so the newest generation on paper is rarely the right one to design into a product today.
- PCIe 1.0 (2003): replaced the parallel PCI and AGP buses with point-to-point serial lanes, the architecture every later generation still builds on.
- PCIe 2.0 (2007): the first doubling, which gave early SSD controllers and higher-bandwidth expansion cards headroom on a single slot.
- PCIe 3.0 (November 2010): moved to the efficient 128b/130b encoding. Still the bedrock of embedded and industrial designs, and the generation we most often qualify in the lab.
- PCIe 4.0 (October 2017): the point where NVMe storage and modern GPUs became bandwidth-comfortable. Today it is the default for demanding new industrial designs.
- PCIe 5.0 (May 2019): the de facto datacenter and workstation standard in 2026, aimed at large-scale data and virtualisation.
- PCIe 6.0 (2022): switched to PAM4 signalling with FLIT encoding for AI inference and training (for example the NVIDIA Jetson platforms) and high-throughput cloud services, often paired with high-performance LPDDR4 / DDR5 memory.
- PCIe 7.0 (specification published by PCI-SIG in June 2025): targets AI accelerators, 800G networking and hyperscale storage. Production silicon is still scarce: this generation is being prepared for, not yet designed in.
For the per-lane and per-slot numbers behind each of these generations, see the PCIe bus speed table above.
These successive PCIe versions guarantee backward compatibility, meaning a PCIe 4.0 or 5.0 device works in a PCIe 3.0 motherboard, but at PCIe 3.0 speeds. This lets organisations evolve their infrastructure progressively without breaking integration with older components.
Why choose AESTECHNO?
- 10+ years of expertise in high-speed electronics design
- 100% success rate on CE/FCC certification
- French electronics design firm based in Montpellier
Article written by Hugues Orgitello, electronics design engineer and founder of AESTECHNO. LinkedIn profile.
PCIe 4.0 vs PCIe 5.0: which generation should you pick?
PCIe 4.0 (16 GT/s, or 2 GB/s per lane) covers the vast majority of industrial designs in 2026: fast NVMe storage, FPGAs, embedded GPU modules. PCIe 5.0 (32 GT/s, 4 GB/s per lane) is only justified when a x4 link has to sustain more than 8 GB/s, typically very-high-rate acquisition or an AI accelerator.
Each generation step is paid for at design time: a halved unit interval, a tighter insertion-loss budget, low-loss PCB laminates, possible retimers and a heavier validation campaign. At AESTECHNO we routinely audit customer PCIe buses with eye-diagram measurements (Tektronix TekExpress suite), and the rule we apply is simple: pick the lowest generation that meets the requirement, and let backward compatibility keep the upgrade path open.
PCIe debugging at bring-up: the link will not come up, or not at the right width
On a new board a PCIe link rarely fails halfway. Either it does not appear at all, or it appears degraded: fewer lanes than expected, or one generation down. Those two symptoms lead to different places.
What the state machine does
Before any data moves, the link goes through the LTSSM, the link training and status state machine. It detects the electrical presence of the partner, negotiates polarity and lane order, agrees a width, then climbs in generation. Each stage can fail separately, and that is what makes diagnosis tractable: the width and generation actually reached tell you where the sequence stopped.
One point that matters for debugging: polarity inversion and lane reversal negotiation are part of the standard. A differential pair swapped in routing does not necessarily stop the link coming up, which hides the error until some less tolerant device exposes it.
Causes, in order of frequency
- Power and reset sequencing. PERST# released before the rails are stable, or too soon after. This is the leading cause of a completely absent link, and it is not electrical in the signal sense at all: it is timing.
- Reference clock. REFCLK missing, out of tolerance, or an inconsistent distribution architecture between the two ends. Without it, no detection.
- Missing coupling capacitors. The standard requires series AC coupling on the pairs. Omitted or misplaced, the link never detects.
- Link comes up at reduced width. Typically x1 instead of x4. One lane is not detecting: an open, a missing capacitor on a single pair, or a partly seated connector. The width obtained tells you how many lanes passed detection.
- Link comes up at a lower generation. That one is signal integrity: loss, vias, length, dielectric quality. At Gen4 and Gen5, equalisation no longer rescues a mediocre channel, it only makes the shortfall show up later in the process.
Telling them apart with what you already own
The first tool is not a scope, it is the host-side link status register. It reports negotiated width and generation, and on Linux lspci -vv exposes both directly, showing maximum supported width alongside the width actually in use. A gap between the two already answers the question.
- Nothing at all: reset, power, clock, coupling capacitors. Check the timing diagram before probing pairs.
- Reduced width: one or more lanes are not detecting. The problem is local, not global.
- Reduced generation: the channel will not carry the rate. Signal integrity.
- Link that trains then drops: insufficient margin or supply noise. The hardest case, and the one that justifies measurement.
What the mistake costs, and when to call someone
A link that trains at Gen3 on a board designed for Gen4 works. That is precisely the danger: the product ships with half the intended bandwidth and nobody notices until load rises. Margin defects are worse still, because they depend on temperature and on the individual unit.
Sorting by width and generation needs nobody and eliminates most cases. What does need support is the degraded generation and the unstable link: the channel has to be characterised, which means simulation before routing rather than observation after fabrication. See our high-speed design page.
Bottom line: 5 PCI Express takeaways
The PCIe bottom line is an operational summary of the five technical levers that decide a successful integration: target generation, stack-up, SI validation, measurement methodology and choice of critical components. The checklist condenses 10+ years of AESTECHNO field experience on PCIe Gen 3 to Gen 5 buses, validated on every recent project against PCI-SIG, JEDEC and IPC reference frameworks. Each of these levers has its own section in this article; the summary below lets you verify in seconds that none was left to chance before committing a layout. An integrator who masters these five points avoids most of the validation failures we see on the PCIe designs that reach our lab.
- Pick the generation against the real loss budget. PCIe Gen 4 (16 GT/s) opens the bandwidth headroom but mandates Megtron 6 or equivalent as soon as cumulative trace length exceeds 12 cm. Gen 5 stays the preserve of datacenter and short links under 10 cm.
- Stack-up and laminate come before routing. We arbitrate Dk, Df, Tg first; high-Tg FR-4 covers Gen 3, Megtron 6 / Isola I-Speed covers nominal Gen 4, Megtron 7 or Tachyon 100G for long Gen 5.
- Systematic back-drilling on Gen 4 and beyond. On our Tektronix TekExpress PCI-SIG test bench, removing the residual via stub through back-drilling restores 20 to 30 % vertical eye opening, with no laminate change required.
- Standardised measurement methodology. Our procedure stands on three steps: TekExpress PCI-SIG compliance, channel insertion-loss / return-loss characterisation on a Keysight ENA VNA, and LTSSM validation under IEC 61000-4-2 / IEC 61000-4-3 EMC stress.
- SI simulation before tape-out. ANSYS SIwave coupled with Cadence Sigrity catches the impedance discontinuities and resonant stubs the Altium DRC lets through. PCI-SIG mask validation belongs in post-layout, never after prototype. NXP and Microchip Switchtec PCIe reference designs publish loss budgets that align with our internal targets.
For supporting context across our high-speed practice, see our USB 3 SuperSpeed versions deep-dive and our complete technical blog.
Conclusion: get the most out of PCI Express
The practical takeaway is straightforward: PCI Express remains the reference solution for interconnecting high-performance components in modern systems. From gaming to AI to datacenter deployments, its high bandwidth, scalable architecture and backward compatibility make it a strategic and durable choice.
AESTECHNO, an expert in integrating and qualifying products with a PCI Express interface, can help you exploit this standard fully and maximise system performance.
Optimise your systems with PCIe. Contact AESTECHNO for tailored PCIe solutions adapted to your technical and industrial requirements.
Contact us to explore PCIe solutions tailored to your projects.
PCI Express: a strategic investment for high-performance systems
PCIe integration is a strategic decision that shapes hardware architecture, the design competencies required and the system's competitive positioning. Understanding the implications lets technical decision-makers balance performance, complexity and time to market.
For CTOs and R&D managers, embedding a PCIe interface in a product is a structural choice that affects hardware architecture, the development budget and the product's competitive positioning. At that level the question is no longer "which generation is best" but "what is the full cost of the target generation", validation tooling included.
Our PCIe expertise: from Gen 3 to Gen 5
At AESTECHNO, our PCIe portfolio spans every current generation up to PCIe Gen 5. We have, for example, designed a custom industrial computer built around an Intel i5 processor with a heavily loaded PCIe architecture - several lanes used in parallel to interconnect NVMe storage, acquisition cards and high-throughput peripherals. That project illustrates exactly what the lane sizing discussed earlier means in practice.
We systematically complement these designs with signal-integrity validation through eye-diagram measurements on PCIe links, to confirm serial-link compliance before industrialisation. Our protocol portfolio also covers DDR2/3/4, LPDDR4, USB 2.0/3.0/3.2 (usb.org), PCIe up to Gen 5, SDI, SPI, I2C, HDMI 2.0, LVDS, MIPI-CSI/DSI, SATA, Bluetooth (bluetooth.com), Wi-Fi, LoRa, RFID, 5G and LTE-M, with RF projects up to 10 GHz.
Field report: PCIe Gen 3 / Gen 4 qualification campaign
On a recent project, in our AESTECHNO lab in Montpellier we measured 18 of 20 PCIe Gen 3 x4 links profiled at 8 GT/s on a Megtron 6 8-layer stack-up. Our measurement methodology stays consistent on every PCIe integration and follows a three-step procedure formalised in-house. Step 1, electrical compliance on a Tektronix bench using the TekExpress PCI-SIG suite: eye-diagram and jitter measurement (Tj, Rj, Dj) on the differential pairs, compared against the PCI-SIG SI / Rx masks published per PCI-SIG. Step 2, channel insertion-loss and return-loss characterisation with a Keysight ENA VNA up to 8 GHz, SOLT calibration performed per the Keysight vendor procedure. Step 3, LTSSM training (Detect, Polling, Configuration, L0) and scrambling validation under EMC stress, measured in a semi-anechoic chamber compliant with the IEC 61000-4-2 and IEC 61000-4-3 series. Contrary to the common assumption that a simple breakout via passes at 8 GT/s without precaution, we found on an 8-layer panel without lambda/10 stitching vias that the return-loss dropped from -22 dB to -8 dB above 4 GHz, a degradation traceable to a return-current path that fragments the reference plane. The field report from the integration team confirmed the fix on the first re-spin: on Gen 4 to Gen 5 ports, the dominant factor is not the laminate but the continuity of the ground reference under each via transition. In our practice across PCIe Gen 3 / Gen 4 engagements, we have observed a recurring pattern: doubling the stitching-via density around differential-pair transitions cuts overshoot by 35 % without touching the stack-up. Despite the schedule tension at the end of the ECO phase, we recommend imposing an SI review with ANSYS SIwave and a parametric sweep up to 6 GHz before Gerber sign-off, a discipline we apply on every customer project. Our internal protocol cross-validates SIwave results with a Cadence Sigrity correlation on critical vias, and Microchip Switchtec PCIe switch reference designs publish similar return-loss budgets for Gen 4 fabric backplanes.
ANSYS SI/PI simulation for PCIe Gen 3/4/5
SI/PI simulation refers to the pre-fab analysis of S-parameters, the eye diagram and the Power Delivery Network on PCIe links at 32 GT/s. The step validates compliance with the PCI-SIG masks before the first tape-out, supported by reference tools such as ANSYS SIwave, Cadence Sigrity and Keysight ADS to close the Gen 5 loss budget.
At AESTECHNO, we systematically simulate PCIe links with ANSYS SIwave and HFSS - Signal Integrity (SI) for the differential pairs at 32 GT/s on Gen 5, Power Integrity (PI) for the Power Delivery Network (PDN) of the root complex and endpoints. In our lab, we have measured that a 0.5 mm via stub at 16 GT/s introduces a resonance inside the useful band that closes the eye diagram by 20 to 30 %, a finding cross-checked with the Tektronix TekExpress PCI-SIG compliance suite. We extract S-parameters, simulate TX/RX eye diagrams with equalisation, and validate the PCI-SIG conformance masks before the first prototype run. On our test bench, this stage has caught more than one via stub or impedance discontinuity that the Altium DRC let through. According to Cadence and according to Keysight, the typical gap between simulation and VNA measurement stays under 1 dB up to 8 GHz when the laminate-model extraction procedure is rigorous, a number we have reproduced on every recent client project. For more on the underlying laminate trade-offs see our high-speed PCB design and PCB stack-up, impedance and EMC guides.
PCB materials for PCIe Gen 4/5
Laminate selection is a Dk / Df / Tg trade-off that becomes critical from PCIe Gen 4 (16 GT/s) onward and structural at Gen 5. On a recent project that combined an Intel SoC with an NVIDIA Jetson module, we observed that switching from high-Tg FR-4 to Megtron 6 cut insertion loss at 8 GHz in half on 12 cm links. On another iteration, with ARM Cortex-A78 cores and a side-car FreeRTOS firmware cross-compiled for a Cortex-M4 MCU, our CI/CD validation ran on GitLab Runners with an automated SIwave simulation pipeline. We are experts in selecting the right material per project: Megtron 6 or 7 for long Gen 5 links (tight loss budget), Isola I-Speed / Tachyon for cost-optimised Gen 4 designs, and high-Tg FR-4 only at Gen 3 or on very short lanes. We balance Dk, Df, Tg, CTE (Coefficient of Thermal Expansion), thermal stability, Pb-free compatibility, manufacturer availability and cost. Our portfolio covers stack-ups up to 28 layers with laser microvias, buried vias and systematic back-drilling on PCIe Gen 5 vias.
PCIe and the embedded-AI market
The explosion of AI at the edge has made PCIe indispensable in any system that needs hardware accelerators. The NVIDIA Jetson platforms and TPU accelerators rely on PCIe as their main interface. At AESTECHNO, we have observed that companies integrating PCIe in their embedded systems gain a real competitive edge in industrial vision, real-time data processing and edge AI markets.
The bandwidth advantage
PCIe bandwidth makes it possible to handle data volumes that simpler buses such as I2C or SPI cannot match. Combined with LPDDR4 / DDR5 memory, a PCIe-based system delivers compute and transfer capabilities that differentiate your product from the competition. That headroom unlocks features (4K video processing, real-time AI inference, multi-sensor acquisition) that used to be the exclusive domain of servers.
When is PCIe justified over simpler buses?
PCIe brings significant design complexity: high-speed routing with controlled impedance, multi-layer PCB stack-ups, and tighter EMC certification constraints. We recommend PCIe when your application demands data rates beyond what SPI or USB can deliver, or when you integrate components that mandate the interface (GPUs, NVMe SSDs, high-performance FPGAs). For more modest needs, an SPI or I2C bus will be a better fit and far cheaper to implement.
FAQ: PCI Express (PCIe)
The FAQ below collects the questions we are asked most often about PCI Express: bus speed, lane configuration differences, version compatibility, hot-plug, and architectural sizing. Each answer draws on what we have seen in the lab.
What is the PCIe bus speed?
PCIe bus speed depends on the generation and on how many lanes the link uses. Per lane it runs 2.5 GT/s on PCIe 1.0, 5 GT/s on 2.0, 8 GT/s on 3.0, 16 GT/s on 4.0, 32 GT/s on 5.0, 64 GT/s on 6.0 and 128 GT/s on 7.0, roughly doubling every generation. Encoding turns that into usable bandwidth of about 2 GB/s per lane on PCIe 4.0 and about 4 GB/s per lane on PCIe 5.0; multiply the per-lane figure for your generation by the lane width to get the real link speed, so a PCIe 4.0 x4 NVMe drive gets about 8 GB/s per direction and a PCIe 5.0 x16 accelerator about 64 GB/s per direction. A quoted generation alone is meaningless without the lane width, and the link always negotiates down to the slowest end.
What is the difference between PCIe x1, x4, x8 and x16?
The number indicates how many lanes (communication channels) are available. Every lane carries data in both directions simultaneously. A PCIe x16 slot offers 16 lanes and is generally used for high-performance graphics cards (up to 32 GB/s with PCIe 4.0 x16). A x1 slot (1 lane) is enough for additional network or USB cards. A x4 card can plug into a x16 slot but will only use 4 lanes.
Is PCIe 4.0 backward compatible with PCIe 3.0?
Yes, PCI Express guarantees full backward compatibility. A PCIe 4.0 card works in a PCIe 3.0 slot (at PCIe 3.0 speed), and conversely a PCIe 3.0 card works in a PCIe 4.0 slot (at PCIe 3.0 speed). The system automatically negotiates the maximum speed supported by the slowest component. This compatibility lets you upgrade progressively without replacing the entire system.
Why move from PCIe 3.0 to PCIe 4.0 or 5.0?
PCIe 4.0 doubles the bandwidth of PCIe 3.0 (2 GB/s vs 1 GB/s per lane), which matters for: high-performance NVMe SSDs (read >7000 MB/s), 4K/8K graphics cards, uncompressed 4K/8K video capture, and AI/ML workloads with multiple GPUs. PCIe 5.0 (4 GB/s per lane) targets datacenters, HPC systems and cloud applications that require massive real-time data transfers.
What is hot-plug in PCI Express?
Hot-plug allows PCIe cards to be added or removed while the system is running, without rebooting. This feature is essential for servers that demand continuous availability (99.999% uptime). Hot-plug requires hardware support (specific slots) and software support (compatible drivers). It is mainly used in datacenters to replace failing Non-Volatile Memory Express (NVMe) cards without interrupting service.
How do I size the number of PCIe lanes I need for my application?
Add up the bandwidth required by every PCIe peripheral. Example: 1 GPU (x16) + 2 NVMe SSDs (x4 each) + 1 10 GbE NIC (x4) = 28 lanes needed. Mainstream desktop CPUs typically expose 16 to 20 lanes, while workstation and server CPUs offer up to 64 to 128 lanes. AESTECHNO can help you size the PCIe architecture against your performance constraints and optimise how the available lanes are distributed.
Related articles
To go further with high-performance architectures:
- I2C bus: operation and applications, a simple multi-peripheral protocol
- SPI bus: integration in embedded systems, synchronous high-speed communication
- High-speed design and signal integrity, PCB design for fast signals
- LPDDR4 memory design, high-performance memory for PCIe systems
- NVIDIA Jetson Orin processors, high-performance embedded systems
- FPGA board design, routing and stack-up for critical signals