NBN upload speeds are rarely the bottleneck for cloud AI use in Australia. Latency is. Sending a 500-word prompt to a cloud AI server takes under one second on any NBN plan, because text is small. Cloud-AI responsiveness depends on the provider, selected endpoint, network route, model, prompt length and service load. US-hosted endpoints add trans-Pacific network latency, while some services offer Australian or regional processing. For multi-turn conversations, that latency compounds. Local time to first token does not depend on the NBN connection when accessed over the local network, but it varies with the model, prompt length, hardware, acceleration, load state and software configuration. That difference in responsiveness is what makes local inference feel faster for conversational use, not bandwidth.
In short: NBN upload speed is not why local AI wins in Australia. Latency to a distant cloud endpoint and privacy can be arguments for local inference, but latency depends on the provider, endpoint region, route, model and workload. Local AI wins for interactive conversation, privacy-sensitive tasks, and offline use. Cloud AI wins for cutting-edge model access, large context windows, and occasional use where hardware investment is not justified.
What Australian NBN Upload Speeds Actually Look Like
NBN plans are marketed by download speed. Upload speeds are lower and, on many plan types, not clearly advertised. The ACCC publishes quarterly broadband performance data showing actual measured speeds across ISPs. The picture for typical Australian households is considerably less impressive than the plan names suggest.
NBN 100 delivers typical upload speeds of 15 to 20 megabits per second on FTTC and FTTP connections. On FTTN , typical upload speeds are 10 to 17 megabits per second and can vary significantly based on line quality and distance to the node. NBN 50 plans typically deliver 17 to 18 megabits per second upload. NBN 25 upload speed depends on the access technology and retail product: current wholesale products include 25/10Mbps and 25/5-10Mbps variants. Fixed Wireless NBN often has asymmetric speeds with limited upload. Starlink is variable from 5 to 20 megabits per second upload depending on load and satellite generation.
| NBN 1000 (FTTP) | Typical upload: 50 to 100 Mbps. Best case for cloud AI prompts. |
|---|---|
| NBN Home Superfast (eligible FTTP/HFC) | Maximum wholesale upload: 50 Mbps. |
| FTTC Home Fast I / eligible FTTP/HFC Home Fast II | FTTC Home Fast I has a maximum wholesale upload of 20 Mbps; eligible FTTP/HFC Home Fast II has a maximum wholesale upload of 50 Mbps. |
| NBN 100 (FTTN) | Typical upload: 10 to 17 Mbps. Variable by line quality. |
| NBN 50 | Typical upload: 17 to 18 Mbps. Fine for text. Slow for large files. |
| NBN 25 | Typical upload: 5 Mbps. Adequate for text. Will feel slow for long documents. |
| NBN Fixed Wireless | Variable: 5 to 25 Mbps upload. Can be congested during peak hours. |
| Starlink | Variable: 5 to 20 Mbps upload. Suited to rural use where NBN is unavailable. |
Why Latency Matters More Than Upload Speed for AI
A 500-word prompt is approximately 3.5 kilobytes of text data. At an upload speed of 10 megabits per second, that transmits in 0.003 seconds. Even at 5 megabits per second on an NBN 25 plan, the same prompt takes 0.006 seconds to upload. For text-based AI use, upload speed is essentially irrelevant. The bottleneck is not getting your prompt to the server.
The bottleneck is where the server is and how long the round-trip takes. Cloud-AI processing locations vary by provider, product and selected endpoint; US and European endpoints are common, but Australian or regional processing is available for some Anthropic and Google deployments. Round-trip network time depends on the user's location, ISP route and selected API endpoint; trans-Pacific endpoints add more latency than nearby Australian or regional endpoints. Every multi-turn conversation exchange includes that overhead. For a simple question-and-answer interaction, 200 milliseconds is not noticeable. For a rapid back-and-forth coding or debugging session with ten or twenty exchanges, the cumulative latency adds up to seconds of waiting that would not exist with a local model.
Streaming responses help but do not eliminate this. Streaming can improve perceived responsiveness, but time to first token includes network transit, service queueing and prompt processing and therefore varies by endpoint, model and workload. Local time to first token varies with model loading, model size, prompt length, context, hardware and acceleration; it should be benchmarked on the proposed system. That immediacy changes the feel of interactive AI use significantly, even when local models generate tokens more slowly overall.
When Local AI Wins in Australia
Local inference running on a NAS or mini-PC serves the local network with round-trip times under 5 milliseconds for wired connections. That responsiveness makes interactive conversation, code assistance, and rapid iteration feel significantly different from cloud AI accessed across the Pacific. For Australian users on any NBN tier, the latency argument for local inference is compelling.
Privacy is the second strong argument. Every prompt sent to a cloud AI service is processed by that provider's infrastructure. For legal documents, medical records, business contracts, customer data, and personal communications, sending that content through a cloud provider is a privacy trade-off that some users and organisations are not willing to make. Local inference processes all data on hardware you control, on your network, and nothing leaves your environment.
Offline use is a third consideration. NBN outages, particularly on FTTN connections in older areas, are a genuine disruption. Local AI inference continues working without internet connectivity. A local model running on a NAS or mini-PC is accessible to all devices on the home or office network regardless of whether the NBN connection is working.
Local AI vs Cloud AI for Australian NBN Users
| Local AI (NAS or mini-PC) | Cloud AI (ChatGPT, Claude, Gemini) | |
|---|---|---|
| First token latency | Varies by endpoint, route, model, prompt, hardware, acceleration and service load | Varies by endpoint, route, model, prompt, hardware, acceleration and service load |
| Upload speed requirement | Not applicable | Trivial. Text is tiny even on NBN 25 |
| Works without internet | Yes, fully offline | No, requires active NBN connection |
| Privacy | All data stays on your hardware | Prompts processed by provider's infrastructure |
| Model capability | Limited by available RAM/VRAM, compute, quantisation and acceptable performance | Access to largest frontier models (GPT-4o, Claude 4, etc.) |
| Context window | Varies by model and hardware; some downloadable models advertise 128K or larger windows, while practical usable context depends on available memory | Up to 1M+ tokens on leading cloud models |
| Cost at high usage | Upfront hardware plus electricity and maintenance; no external API fee per query when fully local | Per-query API cost can scale rapidly |
| Remote access on CGNAT | Requires tunnel (Tailscale or Cloudflare) | Works on any connection without configuration |
| Multi-user household | One device serves everyone on the network | Cost and account requirements depend on the cloud service, subscription terms and aggregate API usage |
When Cloud AI Still Makes Sense Despite NBN
The largest cloud AI models are not available locally. Current closed frontier models from OpenAI, Anthropic and Google are not distributed as downloadable model weights for local execution. They require data centre scale infrastructure. If your use case requires frontier-level model capability, local inference cannot match that regardless of NBN speed.
Context window size is the other hard constraint. Cloud AI models support context windows of 128,000 tokens or more on leading services. On a 16GB system, practical context capacity depends on model size and architecture, quantisation, KV-cache settings and other memory use; some downloadable models advertise 128K windows, but the full window may not be practical on every configuration. Large contexts are also available in downloadable models, but running them locally can require substantial memory and compute, making cloud deployment more practical for some workloads.
For light or occasional use, cloud AI also remains more practical. A household member who uses AI for an hour per week does not benefit from the investment in local AI hardware. Whether cloud AI is cheaper for occasional use depends on the local hardware purchase price, power draw, electricity tariff, utilisation and cloud subscription or API charges. Local AI makes sense when use is frequent enough to justify the hardware.
CGNAT and Remote Access to Your Local AI
Carrier-grade NAT (CGNAT) is a technology used by many Australian ISPs that prevents your home IP address from being directly accessible from the internet. CGNAT is used on some Australian residential broadband services; whether it applies and whether an opt-out is available depends on the provider and plan. CGNAT policy varies by ISP and product; readers should verify the current policy and opt-out options for their exact NBN plan. Aussie Broadband uses CGNAT by default but allows customers to request a no-cost opt-out; static IP is available separately.
CGNAT does not affect local AI inference on your home network. It only matters if you want to access your local AI server from outside your home, such as from a mobile device over 4G or 5G. Without a publicly accessible IP address, you cannot point an external device directly at your home server.
The practical workaround is a VPN tunnel service. Tailscale creates an encrypted mesh network between your devices that works through CGNAT. Your phone running Tailscale can reach your home mini-PC running Ollama as if they were on the same local network, regardless of CGNAT. Cloudflare Tunnel is an alternative that creates an outbound-only connection through Cloudflare's infrastructure, also bypassing CGNAT without requiring an inbound port. Both services have free tiers suitable for personal use.
CGNAT workaround for accessing local AI remotely: Tailscale is the lowest-friction option for most home users. Install Tailscale on your mini-PC or NAS and on your phone or laptop. Your Ollama endpoint becomes accessible as a Tailscale IP address from any location, without any router port-forwarding or static IP requirement. The current Personal plan includes unlimited user devices, up to six users and up to 50 tagged resources. See the Tailscale remote access guide for configuration steps.
What NBN Speed Actually Does Affect
While upload speed does not meaningfully impact typical AI prompt submission, there are scenarios where NBN bandwidth does matter for AI-adjacent workflows. Downloading model files is the most common example. A 7B model at Q4_K_M quantisation is approximately 4 to 5 gigabytes. At an NBN 100 download speed of 70 megabits per second, that downloads in approximately 8 minutes. At NBN 25 speeds (25 megabits per second), the same file takes around 27 minutes. For users on slower NBN tiers in rural or regional areas, downloading large model files requires patience but is a one-time task rather than an ongoing constraint.
Uploading large documents for retrieval-augmented generation (RAG) pipelines is another scenario where upload speed has some relevance. PDF file size varies greatly with scans, images and compression; a 500-kilobyte file would take about 0.8 seconds to transmit at 5 megabits per second before protocol overhead. That is not a meaningful delay even on NBN 25. Multi-gigabyte datasets for local fine-tuning are a different matter, but that is well outside typical home AI use.
For typical text-only prompts, even lower NBN upload tiers generally impose little transmission delay; file-heavy workflows depend on the actual upload tier and file size. The argument for local AI in Australia is not about overcoming NBN limitations. It is about latency, privacy, cost at scale, and offline availability.
Related reading: our NAS buyer's guide, our remote access and VPN guide, and our NAS vs cloud storage comparison.
Free tools: NAS Sizing Wizard and NBN Remote Access Checker. No signup required.
Related reading: our NAS explainer.
Use our free AI Hardware Requirements Calculator to size the hardware you need to run AI locally.
Use our free NBN Plan Finder to compare real NBN plans by upload speed, CGNAT and static IP support.
Does slow NBN upload speed affect ChatGPT or Claude response quality?
No. Text prompts are small enough that even NBN 12 or NBN 25 upload speeds transmit them in milliseconds. Upload speed does not affect the quality of cloud AI responses, only how quickly the prompt reaches the server. The far larger factor is round-trip latency from Australia to US-based AI servers (180 to 250ms), which affects how quickly you receive the first token of a response. Upload speed becomes relevant only if you are sending very large files such as multi-gigabyte datasets, which is rare for typical AI use.
Can I access my home Ollama server from my phone over mobile data?
Yes, but it requires a tunnel if your ISP uses CGNAT. If your residential service uses CGNAT, unsolicited inbound IPv4 access is generally unavailable without an ISP opt-out, public address or tunnel; check your provider and plan. Tailscale is the recommended solution for most home users. Install Tailscale on your NAS or mini-PC running Ollama and on your mobile device. Your Ollama endpoint will be accessible as a Tailscale IP address from any network, including mobile data, without any router configuration. The Tailscale free tier supports this without any ongoing cost.
Which Australian ISPs use CGNAT on NBN?
Many Australian ISPs use CGNAT on at least some residential products, but policies and opt-out options vary by provider and plan. CGNAT policy varies by ISP and product, so users should confirm the current default and opt-out options directly with their provider. Aussie Broadband uses CGNAT by default but allows customers to request a public dynamic IPv4 address by opting out at no additional plan cost. If remote access to a home server is a priority, Aussie Broadband or an ISP that offers static IPs as an included feature is worth considering when choosing a plan.
Is local AI faster than cloud AI on a slow NBN connection?
For interactive conversation, yes. Local AI responds to queries across the local network in under 5 milliseconds round-trip, regardless of your NBN connection. Cloud network latency depends on the selected endpoint and route; a US-hosted endpoint adds trans-Pacific latency, while some services provide Australian or nearby regional processing. The trade-off is that local models generate tokens more slowly once processing begins. Generation speed varies materially with the cloud model and service load and, locally, with model size, quantisation, hardware, acceleration and configuration. Local inference feels more responsive to start but may take longer to finish a long response.
Does using cloud AI on NBN have any privacy risks?
Yes. Every prompt sent to a cloud AI service is transmitted over the internet and processed by the provider's infrastructure. For general queries this is not a concern, but for sensitive content including legal documents, medical records, business contracts, personal communications, or anything containing private data, cloud AI introduces a genuine privacy trade-off. Cloud providers have data retention and usage policies that vary, and jurisdiction questions apply when data is processed offshore. Local AI inference processes everything on hardware you own and control. Nothing leaves your network. Fully local inference can avoid disclosing prompts to a cloud provider, but privacy still depends on secure configuration, access controls, software behaviour, storage and the integrity of the local AI stack.
Setting up local AI inference on a NAS or mini-PC? The local AI hardware comparison covers performance tiers, RAM requirements, and which devices are stocked in Australia.