InfraSense: A Sensing-Aware, Latency-Guaranteed AI Inference Layer for the 5G/6G RAN
A working demonstrator of a dedicated AI inference layer that makes latency a first-class, enforced property of every inference request — together with intent-driven ISAC sensing and SMO-level conflict arbitration. US provisional patent filed.
Dedicated Inference Layer
Latency-Guaranteed Placement — Enforced Tier, Explicit Refusal
The inference layer receives requests comprising sensing data and a latency objective, classifies each request into a latency class, and selects an inference entity from a catalog — subject to a hard tier rule: a latency class admits only infrastructure tiers capable of honoring it. If no compliant placement exists, the request is explicitly refused rather than silently misrouted. Placement further considers live infrastructure metrics, request priority, and an accuracy objective — selecting among model variants (fast / balanced / accurate) to meet accuracy within the latency bound. Each run emits an auditable control action over the interface appropriate to its kind and tier (E2 / O1 / O2).
- Three enforced latency classes — RT (<10ms), Near-RT (10ms–1s), Non-RT (>1s) — with explicit refusal semantics
- Tier-locked placement: RT class admits only DU-adjacent / RT-RIC-adjacent nodes; misrouting is refused, not silently degraded
- Measured evidence: dedicated local GPU inference stayed within 1.3–1.8× of typical latency vs order-of-magnitude swings on shared cloud (50–70s under congestion)
- Validated on two independent local serving engines: vLLM on A100 MIG hardware and a separate engine on H100 — real model calls, not simulation
- Auditable control actions: beamforming, power, slice, handoff, sensing-schedule, or O-Cloud resource recommendation emitted per run
A1 SensingIntent Policies
Intent-Driven ISAC — Declare What the Network Should Sense
A new A1 policy type carries a sensing intent for a geographic/network scope: target metrics (e.g. localization accuracy, sensing coverage), resource constraints (e.g. a ceiling on radio resources or transmit power spent on sensing), a sensing mode, validity time, and priority. An rApp at the Non-RT RIC generates the policy object; the Near-RT RIC derives concrete sensing configuration from its parameters, applying the constraint caps; a closed loop then compares delivered sensing performance against the intent's targets and retunes. The full chain — intent → A1 transmission → derived configuration with caps applied → closed-loop convergence — runs live in the demonstrator using worked examples from the patent filing.
- New A1 policy type: sensing intent with target metrics, resource constraints, sensing mode, validity time, and priority
- rApp at Non-RT RIC generates policy; Near-RT RIC derives concrete sensing configuration with constraint caps applied
- Closed loop: delivered sensing performance compared against intent targets and retuned automatically
- Live demonstrator examples: 0.1m critical tracking intent and 90% coverage intent — both running end-to-end today
- Sits on existing O-RAN A1 interface — adds a policy type, does not replace the stack
SMO Global Agent
Priority-Based Conflict Arbitration Across the Full RAN Estate
A Global Agent at the SMO watches aggregate commitments across sensing intents and rApp/xApp resource consumers. When the sum over-commits the network, the agent resolves by priority: critical intents are protected in full, lower-priority intents are reduced no further than their declared floors, and non-sensing consumers are trimmed proportionally — with the resolution reported for audit. Where no complete resolution exists, infeasibility is reported explicitly rather than papered over. The demonstrator reproduces a 170% over-commitment scenario and resolves it live.
- Watches aggregate commitments across all sensing intents and rApp/xApp consumers at the SMO layer
- Priority resolution: critical intents protected in full; lower-priority intents reduced to declared floors; non-sensing consumers trimmed proportionally
- Live demonstrator: 170% over-commitment scenario reproduced and resolved with auditable output
- Explicit infeasibility reporting — no silent degradation, no papering over unresolvable conflicts
- Operates via O-RAN O1 and O2 interfaces — no proprietary stack replacement required
AODT — Digital Twin
A Living Model of Your Network
The Autonomous Operations Digital Twin (AODT) creates a continuously updated virtual replica of the physical RAN environment. Operators can simulate configuration changes, stress-test edge AI models, and validate new xApps before live deployment — eliminating risk from network updates.
- Real-time synchronisation with live RAN telemetry
- What-if simulation for configuration and capacity planning
- xApp validation sandbox before production rollout
- NVIDIA Omniverse integration for 3D city-scale visualisation
Artifact Catalog
Curated AI Models for Every Use Case
The OranSense Artifact Catalog is a governed repository of pre-trained sensing models, xApps, and inference pipelines — ready to deploy on any O-RAN-compliant infrastructure. Operators and enterprise customers can browse, license, and deploy in minutes.
- Pre-trained models for crowd analytics, health monitoring, and security
- Versioned xApp packages with automated compatibility checks
- One-click deployment to OranSense edge nodes
- Partner-contributed models via the OranSense Marketplace
Ready to explore the platform?
Request a technical briefing or live demo with our engineering team.
Request a Demo