Why Singapore’s large language model isn’t sweating GPT-4
WHAT does the late Indonesian President Suharto, who ruled the country for 31 years, have to do with generative AI and the need for locally developed large language models (LLMs)?
Citing the long-deceased leader as an example, AI Singapore (AISG), a nonprofit organisation connecting the country’s AI research institutions, startups, and companies, argues that Western-developed LLMs may harbour certain biases stemming from their training data and cultural context.
For instance, a version of Llama 2, the open-source LLM developed by Meta, portrayed Suharto largely negatively, emphasising his role in political suppression and human rights abuses.
At a recent press event, AISG demonstrated how its model Sea-Lion, which stands for South-east Asian Languages In One Network, also highlighted the achievements of Suharto, such as creating an environment of political stability and economic prosperity.
This suggests Sea-Lion’s potential for handling nuanced perspectives on sensitive topics, particularly those requiring local context.
US-China AI arms race
One key argument driving organisations like AISG to develop their own LLMs is the need for models that are not solely aligned with Western culture.
Last month, Singapore announced its revised national AI strategy and committed S$70 million to further develop Sea-Lion. However, some AI practitioners question the practicality and timing of the LLM project in light of OpenAI and other tech giants’ rapid advancements.
Still, amid the escalating US-China AI arms race, AISG believes there is a need for diverse LLM alternatives.
Leslie Teo, a senior director at AISG, pointed out that perceived biases in AI models do not stem from malicious intent or personal values of their creators. Instead, the origin and linguistic scope of training data play a crucial role in shaping model bias.
According to AISG, which collated data provided by AI tool developer Hugging Face, about 73 per cent of existing LLMs originate from the US and China, and 95 per cent of all models are primarily trained on English-language data or on a mix of English and one of Chinese, Arabic or Japanese.
This means South-east Asian languages such as Bahasa Indonesian, Thai, and Tamil are severely underrepresented. As an example, AISG says that less than 0.5 per cent of Llama 2’s training data comes from the region’s languages.
Sea-Lion aims to bridge this gap, claiming to be the first open-sourced LLM specifically focused on South-east Asian languages and contexts.
Around 64 per cent of the data in the model is in English, while South-east Asian languages comprise 13 per cent, with the remainder taken up by Chinese and code.
It was trained on one trillion tokens of data and comes in two variants – three billion and seven billion parameters. Generally, models that are trained on a higher number of parameters will be more powerful.
For example, the largest version of Llama two has 70 billion parameters, while GPT-4 is said to have 1.7 trillion.
The team behind Sea-Lion is 70 per cent Singaporean, with other members coming from countries including Indonesia and Vietnam.
To build Sea-Lion, they had to overcome the scarcity of high-quality, publicly available data in South-east Asian languages. Making the task more complex, Sea-Lion was built from scratch and has been trained only on copyright-free data to avoid claims of copyright infringement. Data that is subject to copyright protection tends to be of higher quality.
Building a model from scratch, however, allows the developers to curate and clean the training data, ensuring that it accurately reflects the diverse linguistic landscape of South-east Asia.
This approach also means that companies can be more confident in using Sea-Lion’s output because it will not infringe on copyright laws.
“If you didn’t do it from scratch, you don’t have full control over the data that you use,” said Teo.
Indeed, the issue of AI models using copyrighted data has been a growing concern. The New York Times’ lawsuit against OpenAI, alleging that the latter copied millions of articles from the newspaper to train its LLMs, exemplifies how content creators are increasingly pushing back against this trend.
Lim Swee Kiat, co-founder of AI image generation app Pebblely, said that larger clients are “getting very concerned” about where training data for LLMs is sourced, suggesting growing anxieties surrounding AI-generated content and its potential copyright implications.
Battle of the models
Sea-Lion was also tested on questions beyond Suharto, including identifying the current president of Indonesia and why the country plans to relocate its capital city.
A total of eight questions were posed to the model – five in Bahasa Indonesian, two in Thai, and one in Tamil.
Compared to Meta’s Llama 2, Alibaba’s SeaLLM, and OpenAI’s GPT-4, Sea-Lion appeared to outperform its competitors in terms of speed, accuracy, and succinctness. However, GPT-4 compensated for its slower response time with more context-laden answers, which were also accurate.
Meanwhile, unlike the other three models, Llama two responded in English to questions in Bahasa Indonesian, while the quality of SeaLLM’s responses was patchy, with some answers exhibiting hallucinations. For example, the question in Tamil resulted in gibberish.
AISG expects Sea-Lion’s performance to be even better when queries are made in non-Latin languages such as Tamil and Thai, which are typically underrepresented in training data for existing models.
That said, it is important to remember that this limited demonstration may not fully represent the models’ overall capabilities.
AISG has proposed a new benchmark called BHASA to evaluate the performance of LLMs in South-east Asian languages.
It said a new benchmark is necessary, as most existing ones only measure LLM performance in English.
BHASA consists of three criteria: first, how effectively the model understands, reasons, and generates text in a specific language; second, how well the model adheres to traditional linguistic measures such as syntax and semantics within that language; and third, how adept the model is at interpreting the cultural context of that language.
Using this benchmark, the AISG team found that Sea-Lion is placed only second to GPT-4.
‘Not trying to beat GPT-4’
Unlike GPT-4, Sea-Lion is free and fully open source, which means that users can download both the model itself and the training data, the official adds. On the other hand, OpenAI does not disclose the training data it used for GPT-4.
That said, GPT-4 has an API, which makes for easier access by third parties. AISG said that it currently has “no plans to offer an API” since it has not built a chatbot and does not want its model to be misused.
Though it might sound too good to be true, AISG reiterated that using the model only requires adhering to its open-source licence.
Indonesian ecommerce platform Tokopedia is one of the earliest companies testing the model for business purposes – it is using Sea-Lion to generate product descriptions in other South-east Asian languages. Another company is using the LLM to build a chatbot that provides legal advice in Bahasa Indonesian, Teo added.
Given that Sea-Lion is not a commercial endeavor, what will success look like?
Teo expressed enthusiasm for the potential of Sea-Lion’s data to be used in other models, for example, when they are presented with queries in South-east Asian languages.
He also emphasised AISG’s role as a catalyst in Singapore’s AI ecosystem. The goal is to cultivate a robust local talent pool equipped with expertise in all aspects of genAI.
Teo expects companies and startups using Sea-Lion to create jobs in Singapore, “jobs that would otherwise be in Shenzhen or Silicon Valley”.
“We are trying to build public infrastructure that is necessary in the AI space so that lots of people will use AI,” he added.
Given the potential benefits of Sea-Lion, the US$52 million allocated by the Singapore government for LLM research and development might seem insufficient.
Teo, however, has taken a different perspective, stating that “too much money can be bad.” He believes that budget constraints can drive innovation.
Despite its nonprofit status, AISG must still demonstrate its value proposition. It is still early days for Sea-Lion, and its success hinges on finding product-market fit amid the ongoing global genAI race, where billions of US dollars continue fuelling the competition. TECH IN ASIA
TRENDING NOW
Despite the de-dollarisation debate, demand for dollar liquidity in Asia is growing
URA to review guidelines on floor space to give developers more design flexibility: Chee Hong Tat
He built the Vingroup empire. Now South-east Asia’s richest man is handing some key roles to his sons
Can a first-time homebuyer couple earning S$18,000 a month afford a new EC unit?