Choosing the right model for a specific task is one of the problems we face when implementing LLM based projects or basically any RAG systems
So some cool websites I found for that is https://artificialanalysis.ai
It is great to check the latency (the time that u spent before receiving the first chunk of output), the price, the output speed and so on.
the other site is lmarena
https://arena.ai
it is kind of crowdsourcing scores of language models by Humans feedback
For normal dev u can do plug and play but fro production u should design ur system so that u can use 2 or more models and ultimately compare the user satisfaction or some other metrics u get since it's catchy to only trust the benchmarks and leader boards only.
Other than that u have studios like google ai studio, groq ai studio and so on that u can try out the models right there with some limit. U can also adjust the temperature.
for choosing between local models for those who have gpus use lmstudio https://lmstudio.ai/
#resources@qalabit
So some cool websites I found for that is https://artificialanalysis.ai
It is great to check the latency (the time that u spent before receiving the first chunk of output), the price, the output speed and so on.
the other site is lmarena
https://arena.ai
it is kind of crowdsourcing scores of language models by Humans feedback
For normal dev u can do plug and play but fro production u should design ur system so that u can use 2 or more models and ultimately compare the user satisfaction or some other metrics u get since it's catchy to only trust the benchmarks and leader boards only.
Other than that u have studios like google ai studio, groq ai studio and so on that u can try out the models right there with some limit. U can also adjust the temperature.
for choosing between local models for those who have gpus use lmstudio https://lmstudio.ai/
#resources@qalabit