What external GPUs work reliably in the x4 slot?

I’ve been incrementally researching the Framework Desktop for AI inference purposes, and have asked a couple of questions here, and have been going through a few conversations with AI. Slowly and iteratively, my understanding (of local AI and the Desktop’s capabilities) improves.

My hope is that while I will subscribe to commercial AI providers (for code generation) for now, I would like to run my AI capability entirely myself in the future. I am pretty much sold on the Desktop, but would like to make this a future-proof decision. I hope to be able to use the onboard unified arch for my AI workloads, but if that does not suffice, I’d like to be able to keep the eGPU option in my back pocket.

Thus, I have been fascinated to see that people are buying the motherboard on its own, and putting it in a custom case. This of course permits a solid GPU to be used if the unified memory arch proves insufficient for a given compute task. It should be noted also that the standard case does not have space for a GPU, nor an easy way to route cables to an external housing.

However being new to GPU cards (and not being a gamer) I am struggling to understand what cards will fit. It looks like lower-spec ones are OK, but higher-spec ones will need x8 or x16. I’ve looked at some card sellers though, and their specs tend not to mention what a card needs (maybe experienced buyers can infer it through spec details I do not understand). Does Framework have a list, or is there any such thing in this forum?

The desktop electrically may only have a x4 slot but you can buy riser cables that can turn that into a physical x16 slot. It will still only work at x4 PCIe lane speeds but any card will do. Especially if you are willing to mount the motherboard in an alternate case. If you prefer to not use a riser cable, you can break the very end of the slot off to make it open-ended and then insert whatever card you desire. Again, it will only run at x4 speeds but it will work fine, especially for AI inference tasks where bandwidth isn’t as necessary. Personally, I’d just do the riser cable tho.

If you are asking what cards will natively fit in that port then be prepared to spend an absurd amount of money.

EDIT: No card will natively fit as the PCIe slot isn’t exposed at the back. So you have to mount it into a normal case. So just use a riser cable and call it a day.

EDIT 2: I’d buy this as I know that company also makes well regarded Thunderbolt eGPU accessories so I’d trust it actually will do what it says it’s supposed to do. I’d get the shortest length you can manage for signal integrity concerns.

EDIT 3: If you just want a list of the lane count that any GPU is built with, just use the TechPowerUp GPU database. Just keep in mind that for AI workloads, lane count doesn’t matter too much for what you are doing.

1 Like

i tried a pcie 4 x4 to oculink adaptor to connect an rx 7900 xtx via minisforum deg1, but couldn’t make it work. the only time the gpu was detected was when i changed the pcie bus speed to gen 3. however, the desktop (fedora silverblue) was extremely sluggish with that config and locked up frequently (so impractical).

i’m currently in the process of returning all the gear related to the egpu. i might give this setup a try at a later point because it’s extremely valuable. nonethesless, the desktop igpu is powerful enough to run local ai productively.

note that the framework case doesn’t expose the pcie x4 slot. so you’ll want to get your own case. although i couldn’t make the egpu work, i’m extremely happy with the jonsbo z20 orange case.

it seems such a missed opportunity to have a well documented, fully supported path to egpu via oculink.

1 Like

Why made you go the eGPU route rather than the riser cable or other direct connection?

the z20 is structured such that if the pcie x4 slot on the board was an x16 slot instead, the gpu would just plug right in, and be correctly exposed through the back of the case. there would be no need for a riser cable (unlike the terra case, for example). i tried using this, this and this x4 to x16 adaptor, but then there wasn’t enough clearance for the gpu.

i suppose using a bigger case or a smaller gpu could have helped. i just didn’t have any more energy left to keep trying alternatives after a last ditch microcenter visit didn’t help either.

1 Like

Super, thanks. I like the idea of this, though this search doesn’t bode too well!

That accords with some AI advice, which is a relief given that the GPU Database doesn’t reassure. It said these are good:

Nvidia RTX 3050 x4 Single-slot, works great.
Nvidia RTX 3060 Ti x8 (works in x4) Mid-range, good for AI.
AMD RX 6400 x4 Single-slot, low power.
AMD RX 6700 XT x8 (works in x4) High VRAM, good for AI.
Nvidia RTX 4060 x8 (works in x4) Newer, efficient.

It said I should avoid:

  • RTX 4070/4080 (often need x8/x16).
  • RX 7800 XT/7900 XT (often need x8).

At another time it gave me this:

GPU Model Likely Works in x4? Notes
Nvidia RTX 4070 :warning: Maybe May work but could bottleneck.
AMD RX 7800 XT :warning: Maybe May work but could throttle.
Nvidia RTX 4080 :cross_mark: No Likely to fail or overheat.

I have no idea how to verify the robot claims though - could all be hallucinations :ghost:

The issue is that many cards might electrically only have 4 lanes of bandwidth but they have the same physical connector as a card that uses all 16 lanes. Hence the need for the riser cable. You don’t need to worry about any card “failing” or “overheating” as the AI said. Crypto miners used to do this all the time as mining crypto (much like AI) is VRAM intensive, not bandwidth intensive.

Get a GPU with at least 16GBs of VRAM but 20 or more would be better. AMD if you plan to use Linux ever. Faster cards will return results faster but the bigger memory pool will ensure you can use the larger models. Ignore low end cards like 3050 and 6400, they’re worthless for AI.

EDIT: When configuring the FW Desktop to buy, make sure you get the biggest amount of RAM you can afford. It can’t be upgraded after unless you’re handy with a reflow oven and reballing memory chips. That RAM is shared between GPU and CPU and the more you have means the bigger the model you can stuff in there and that matters more than how fast the GPU core is. I wouldn’t buy a dGPU until you at least try out how the iGPU performs and see if it works for your needs.

1 Like

Many thanks for your advice. Yes, I’d be getting the 128GB Desktop as the base. I am somewhat minded to put the mobo in a custom case so that I can add a GPU card, but defer the purchase of that card until a later time. Of course I could just get the standard case and then sell the extraneous pieces if I do the upgrade. Decisions, decisions…

It is a relief that more substantial cards are still open to me - the AI getting that so substantially wrong shows that this arena is perhaps still in need of experienced, human advice :face_with_tongue:

1 Like

FWIW I got local inference running the past few days, have Nvidia parakeet 0.6b set up for speech to text (pretty fast, I’m satisfied) and have Qwen 3.6 35b A3B running too. My machine is crippled by a janky cooler so only short bursts are representative of its normal working output, but my impression on the tokens/second measured on a short burst is if you want high thinking, this 8060 may not cut it…

I’m comparing it to using codex with 5.5 low, it’s just far far faster and also smarter.

As far as dedicated GPU goes, I don’t think it makes sense to buy something this expensive to try adding a dedicated GPU. It makes more sense to buy a cheaper CPU and spend more on the GPU. You can also go up to DGX Spark for more FLOPS coupled with 128GB memory, that should run LLMs faster than this.

I got a strix halo because I can use 64gb of ram with lazily written python, the GPU is just a bonus. It’s pretty capable, but I expect subscription models to be doing the heavier lifting.

My hope, per earlier comments, is that I will be able to get rid of them in a few months time :tada:

I see what you mean. My thought process is that I’m very keen to support the FW philosophy, and it is quite possible that Strix Halo might meet my AI needs. However, if it does not, I would then have a base machine on which a GPU card can be installed (though I note that the four-lane arch of the PCI interface may not be ideal). I’ve flip-flopped several times on increased VRAM (Strix) versus better performance (Nvidia) and I’m coming down on the side of FW/Strix for now.

I’ve seen some comments in LocalLLaMA on Reddit that some of the newer code-writing models fit into the VRAM of a consumer-grade GPU, negating the VRAM benefits of the unified arch approach. However, I’d like to try both; I am estimating that each AI user has a specific view about what output “feels right”, and there isn’t much of an alternative to trying it for oneself.

I acknowledge I’d be hedging my bets, but in the worst-case scenario, the unified arch does not perform well enough for me, I sell the unit nearly-as-new, and I move the GPU to a cheaper (if less repairable) machine.

FWIW my friend at OpenAI assured me the data protection is serious, they have lots of big corporate customers. Honestly, the subscription prices are fantastic and they probably lose a lot of money hosting them. It’s good to have a backup option but I think my productivity would fall a lot.

I do plan on doing more testing to see what these little local free models can do, but the productivity difference is pretty large if you’re not already very familiar with what you need to write.

Do you already have a smaller model that fits on 1 or 2 dedicated GPUs that you think is good enough? The biggest problem is getting a lot of GDDR6 is much more expensive than DDR5, e.g. >3k for a modded 4090 48gb. V100s are cheap for the memory but the actual processor doesn’t have native BF support. I think something like 3x 3080s or 2090s is probably the cheapest way you can get >60gb VRAM and that’s still somewhat small in the face of >200b models.

I don’t doubt this view is widely held. My view is that the cloud AI providers are all ethically compromised (the sub-topics for which being covered very well elsewhere) and once my brain has taken a view, there’s no use you (or me) trying to persuade it otherwise :zany_face:. I assure you it is very stubborn indeed! :brain:

I am also pondering whether my AI coding style will be based on sequences of atomic commits, or solid-planning plus several hours of computation resulting in a “one-shot” solution. Currently atomic commits is winning, and if my complexity suspicions are correct, this requires much less AI horsepower, and thus non-frontier models may work well for me.

I do not. I’ve played around with some small models running in CPU-only mode, and I’ve the first tier Claude subscription. So there is plenty of ground for me to investigate! I don’t have any dedicated hardware for this project as yet, which is the reason for much of my recent posting on this forum.

FWIW my GPU intentions are probably closer to @GhostLegion’s post above - in the order of 24GB and likely not above 32GB. My budget is in the order of £3-5k, I think, and if I do get a GPU card, I shall not be averse to buying second-hand.

My view is that the cloud AI providers are all ethically compromised (the sub-topics for which being covered very well elsewhere) and once my brain has taken a view, there’s no use you (or me) trying to persuade it otherwise :zany_face:.

Don’t get me wrong, I probably hate them more than you… but productivity is productivity. For now, they’re giving you a state of the art tool for a subsidized price.

If they remove the subsidies and then nerf the products, the next best thing would probably be to use one of the lighter models hosted by a cloud provider, but those aren’t being subsidized. I looked into renting a server, 1 year costs the same as buying the actual hardware yourself. Buying the hardware…is a s***load of money.

There’s just no way around it, they have 11 digits of capital and you don’t :frowning:

1 Like

I can send that advice onto my :brain: but I am pretty sure I already know what the answer :pouting_cat: will be! :laughing:

My research thus far is that AI code gen (or my modest requirements for the same) may not require the horsepower you’ve alluded to. But it shall be a fun project to answer that question. :nerd_face:

That’s already somewhat starting with the gradual shift to per token billing. Both OpenAI and Anthropic are not profitable and are looking to IPO. Revenues must come up to look better to investors. I personally don’t think this bubble can last even 1 more year. I do believe AI can change the world but only when capable models can be run locally on reasonably priced hardware. Not necessarily phones but at least laptops.

1 Like

Eh, the logic behind putting hardware on the cloud is that expensive hardware can have higher utilization rate. The equilibrium is prices converge to the cost of cloud compute plus a thin margin. Cloud compute itself should also see margins thin. Unless your computer is being pushed a significant number of hours during the day, it will likely be more expensive than cloud compute. And if anyone’s software has an edge over the competition, they wouldn’t give it out for free.

I see what we can buy now for $20/month as a gift that may not last much longer. I hope my 395 can fill in the gaps if it gets more expensive in the future, but I need the RAM anyways, I burn through a lot using lazily written python.

That’s true for literally most PCs for most tasks, what’s your point? PC’s race to idle and it’s why idle power draw is so important. Doesn’t mean I’m willing to subscribe to GeForce Now or a VM to compose a Word doc. I’m not suggesting every needs a Xeon, I’m saying when AI can run well on whatever we call a Core i5 is when it’ll be most useful. Or midrange GPUs. Or NPUs. Whatever the hardware choice is.

For AI? Yes, absolutely. For now. It won’t last and it can’t last.

1 Like

That’s Office 365, right? :laughing:

1 Like

It’s called inequality. You need a hundred times more money than the average PC to run a frontier model. If you need the marginal increase in capability for some business need, good luck competing with a normal consumer device.

AI makes intelligence more and more about how much money you have.

I find the discussions about AI to be fascinating because they’re as much about social theory, economic models, personal political beliefs etc. as they are about technology. I am cautious in terms of adoption, strident in terms of ethics, and perhaps optimistic about (eventual) local running. With my philosophical hat on, I wonder if I didn’t make these choices anyway…

With this morass of intersecting topics, there are invariably going to be areas where we don’t find agreement, and that’s not just OK, that’s inevitable. :relieved_face: