TheMachine Press

A daily newspaper for the age of artificial intelligence.

Morning editionPermanent story

chips infrastructure

The Edge Chip Stopped Choosing Between a Pipeline and Parallel Work

Para-Pipe maps operator concurrency within and across stages, producing Pareto choices for latency, throughput and energy on heterogeneous SoCs.

Published Updated Story ID: mp-2026-09-06-013
Read the complete editionStory JSON

Summary

Para-Pipe maps operator concurrency within and across stages, producing Pareto choices for latency, throughput and energy on heterogeneous SoCs.

The framework searches how a neural graph should share work across big and little CPU cores, a GPU, DSPs and a dedicated accelerator. On one Amlogic system, throughput-optimized configurations improved average energy efficiency by 11.0 percent over pure pipelining and 23.3 percent over non-pipelined parallel execution. A second automotive-class platform supplied another heterogeneous test. These are measurements on the authors' graphs and devices, not general efficiency guarantees for all edge workloads.

Why it matters

Para-Pipe maps operator concurrency within and across stages, producing Pareto choices for latency, throughput and energy on heterogeneous SoCs.

Limits and context

  • These are measurements on the authors' graphs and devices, not general efficiency guarantees for all edge workloads.

Key claims

  1. Para-Pipe maps operator concurrency within and across stages, producing Pareto choices for latency, throughput and energy on heterogeneous SoCs.

    Qualification: These are measurements on the authors' graphs and devices, not general efficiency guarantees for all edge workloads.

    Evidence: source-2026-09-06-013

Sources

  1. arXiv preprint 2609.04168arXiv · primary research

Corrections

No corrections have been recorded for this story.