Faster AI chat model

janubaba

Junior Member
Joined
Jun 15, 2023
Messages
106
Reaction score
34
Hi

are you u sing all available AI chat models from open ai and bing and google? what do you think about the speed of response? which you think is most faster? chatgpt from openai is mostly slow.

please share you experience for other chat models.
 
Computation time follows a linear correlation to the number of parameters of said model. (Assuming equivalent hardware)

That being said, out of Llama 3 (70 B) and GPT-3 (175 B), Llama 3 would on average be more than twice as fast than the latter. Of course it should be tested experimentally for your use-case.

But regarding the speed of the response from clicking the button to receiving the message, you're now talking about network latency and bandwidth, which is is the main bottleneck here.

This is also highly subjective because it depends on your location from the server. I'd focus more on which model produces the result you're looking for over which website is faster.
 
google is faster compare to others. also i never used bing ai
 
In terms of response, I found 4o mini to be faster than 4o. Quality wise 4o is better.
 
The nost popular chatgpt is fastest I think so
 
In terms of speed, groq.com is probably fastest and it runs open-source models only. They have specialized hardware and they claim to be the fastest in AI inference. Whatever.. they have a free tier so you can always try for yourself.
As for the big guns, Gemini from Google is also fast, at least compared to ChatGPT free. The quality is similar to open models but below ChatGPT on certain tasks. Sometimes better, sometimes worse, it depends on your use case.
If you only need speed, run a TinyLlama or other similar small model on vast.ai . But quality is submediocre.
 
I think this is based on preference , I mean speed can be variable, and it’s a good idea to test different models to see which one best fits your needs.
 
I feel them about the same, maybe you use it when they are having more traffic and get slow but so far I feel all of them are about the same speed don't see difference but also could be what I suse them for, just general and simple information.
 
Look into Groq. They have near instant inference speed for <10B parameter models with their specialized hardware. 10x inference speeds across the board compared to the competition. There are other labs working on it but nothing much as easily accessible.
 
Dall-E is the best for images. Also Chatgpt is good for blog posts
 
Back
Top