gemini-2.0-flash-exp vs llama-3.3-70b-instruct

Pricing, Performance & Features Comparison

gemini-2.0-flash-exp

Authorgoogle

Context Length1M

Reasoning

Providers1

ReleasedDec 2024

Knowledge CutoffAug 2024

License-

Gemini 2.0 Flash-Exp is an experimental low-latency model from Google Vertex AI that supports multimodal inputs and outputs, including real-time vision, audio streaming, and text-to-speech. It provides improved performance over earlier Gemini releases, offering features such as bounding box detection, native image generation, and complex function calling. The model excels at agentic tasks and is suitable for scenarios requiring fast responses and versatile tool use.

Input$0.00

Output$0.00

Latency (p50)-

Output Limit8K

Function Calling

JSON Mode

InputText, Image, Audio, Video

OutputText, Image, Audio

google-vertex

in$0.00out$0.00--

llama-3.3-70b-instruct

Authormeta

Context Length128K

Reasoning

Providers1

ReleasedDec 2024

Knowledge CutoffDec 2023

License-

Llama 3.3 is a text-only 70B instruction-tuned model that provides enhanced performance relative to Llama 3.1 70B–and to Llama 3.2 90B when used for text-only applications. Moreover, for some applications, Llama 3.3 70B approaches the performance of Llama 3.1 405B.

Input$0.45

Output$0.45

Latency (p50)-

Output Limit4K

Function Calling

JSON Mode

InputText

OutputText

avian

in$0.45out$0.45--