The semiconductor and AI stack is moving into a new phase shaped by efficiency, scale, and control.
This week shows falling AI inference costs, new ways to reduce memory usage, and continued investment across photonics, EUV, and advanced materials.
At the same time, governments and hyperscalers are strengthening local supply chains and building their own AI capabilities.
From power systems to AI platforms, the focus is shifting toward tightly integrated systems where compute, energy, and data all need to work together.
1) Gartner Sees AI Inference Costs Falling Sharply by 2030
Efficiency gains reshape token economics
Gartner expects the cost of running inference on large language models to drop by more than 90% by 2030, driven by better chips, higher utilization, and more efficient model design.
At the same time, performance could improve up to 100× compared to early 2022 systems, showing how quickly the full stack is evolving.
Lower cost per token may not reduce total spend
Even with cheaper inference, total costs may still rise. More advanced AI systems use significantly more tokens per task and run more complex workflows.
This is why model orchestration is becoming critical, with smaller models handling routine work and larger models used only where needed.
2) NVIDIA Positions AI as a Mix of Open and Proprietary Models
AI moves toward a multi-model ecosystem
At GTC, NVIDIA outlined a future where open and proprietary models work together rather than compete directly.
The company introduced the Nemotron initiative, bringing together partners like Mistral AI to build open foundation models with shared compute, data, and expertise.
Orchestration becomes the key layer
The focus is shifting to how models are selected and combined. Different models will handle different tasks based on cost, performance, and domain fit.
This makes orchestration software one of the most important layers in the AI stack.
3) Google Targets AI Memory Bottlenecks With TurboQuant
New approach to compress model memory
Google introduced TurboQuant, a method to compress key-value cache data down to 3 bits while maintaining accuracy.
This directly targets one of the biggest bottlenecks in AI systems: memory usage during inference.
Impact on hardware demand is unclear
While the approach improves efficiency and speeds up computation, it could reduce memory demand per workload.
At the same time, lower costs may increase overall usage, meaning total demand for memory could still grow.
4) Lumentum Expands U.S. Photonics Manufacturing
New facility supports AI data center demand
Lumentum is building a new indium phosphide manufacturing site in North Carolina to produce lasers for AI data centers.
The facility is expected to ramp in 2028 and will support large-scale optical networking needs.
Onshoring aligns with AI infrastructure buildout
The move strengthens domestic supply of critical optical components, with customers including NVIDIA.
It reflects a broader shift toward securing photonics supply chains closer to AI infrastructure deployment.
5) AIXTRON Expands Manufacturing Footprint in Asia
Malaysia site adds regional capacity
AIXTRON will build a new facility in Malaysia focused on assembly and testing for compound semiconductor equipment.
The site supports technologies used in power electronics, lasers, and communications.
European core remains unchanged
The expansion improves access to Asia’s supplier ecosystem while keeping R&D and core production in Europe.
This shows how equipment players are balancing global supply chains with regional strengths.
6) Infineon and DG Matrix Target Data Center Power Bottlenecks
Silicon carbide enables more efficient power conversion
Infineon Technologies and DG Matrix are working on solid-state transformers using silicon carbide.
These systems can be smaller, lighter, and more efficient than traditional transformers.
Power becomes a key part of AI infrastructure
As AI data centers scale, power conversion is becoming just as important as compute and networking.
This creates a new growth area for semiconductor companies focused on power devices.
7) Arm Moves Into Server Silicon for AI Workloads
From IP to full chip design
Arm Holdings is launching its first production CPU designed for AI data centers.
The new processor targets agentic AI workloads that require constant orchestration and data movement.
Backed by major ecosystem players
The platform is being developed with partners like Meta and supported by a wide range of cloud and system companies.
This move brings Arm closer to competing directly at the silicon level in AI infrastructure.
8) Google Expands Quantum Strategy and Sets 2029 Timeline
Dual approach to quantum hardware
Google is developing both superconducting and neutral atom quantum systems to balance performance and scalability.
This increases flexibility in its long-term quantum roadmap.
Push toward post-quantum security
Google also set a 2029 target for transitioning to post-quantum cryptography.
This reflects growing concern around future risks to current encryption systems.
9) Apple Expands U.S. Semiconductor Supply Chain
New partners added to manufacturing program
Apple is expanding its domestic supply chain with partners like Bosch, Cirrus Logic, and TDK.
The focus is on sensors, mixed-signal chips, and advanced materials.
Part of a broader investment strategy
The initiative is linked to a $600 billion long-term investment plan, aimed at strengthening local production and supply resilience.
10) Lace Develops Alternative to EUV Lithography
Helium beam approach targets atomic-scale patterning
Startup Lace is developing helium atom beam lithography, which could enable much smaller features than current EUV systems.
This could extend scaling beyond the limits of existing photolithography.
Still early, but strategically important
The technology is in early stages, with a pilot fab planned around 2029.
If successful, it could become a long-term complement or alternative to EUV.
11) Air Liquide Expands Materials Production in Taiwan
New plant supports advanced chip manufacturing
Air Liquide opened a new facility in Taiwan focused on advanced materials used in deposition and etching.
These materials are critical for atomic-scale semiconductor processes.
Closer integration with foundries
Producing materials locally helps speed up process development and improves yield during production ramp.
12) Rigetti Plans 1,000-Qubit Quantum System in the UK
Major investment in scaling quantum systems
Rigetti Computing plans to invest up to $100 million to build a system with more than 1,000 qubits in the UK.
The project supports the country’s broader quantum strategy.
Focus on hybrid systems
The approach combines quantum processors with classical computing, which remains key for practical applications.
13) South Korea Invests in AI Chip Startup Rebellions
Government backs domestic AI silicon
South Korea is investing $166 million into Rebellions to support development and production of AI accelerators.
The move is part of a broader push to reduce reliance on foreign chip suppliers.
Focus on building a local ecosystem
The investment highlights growing national efforts to build independent AI hardware capabilities.
14) Alibaba Pushes Token-Based AI Business Model
Shift from compute to value-based pricing
Alibaba Group is focusing on token-based economics and agentic AI systems.
The goal is to move beyond selling compute and toward delivering full AI-driven workflows.
MaaS becomes a core revenue driver
Machine-as-a-Service is expected to become a major growth engine, supported by increasing adoption across enterprise and consumer platforms.
15) SK hynix Expands EUV Capacity for AI Memory
$8 billion investment in EUV systems
SK hynix is investing heavily in EUV lithography to scale next-generation DRAM and HBM.
This supports growing demand for high-bandwidth memory in AI systems.
Memory becomes a key AI bottleneck
As AI workloads grow, memory performance and capacity are becoming just as critical as compute.
The investment reflects rising competition in advanced memory technologies.