Libove Blog

Personal Blog about anything - mostly programming, cooking and random thoughts

Missed Shot at Artificial General Intelligence

The stated goal of OpenAI is "to ensure that artificial general intelligence benefits all of humanity". With the release of ChatGPT, they might have missed humanity’s only shot at creating an Artificial General Intelligence (AGI).

It's all about data

The performance ("intelligence") of large language models (LLMs) mainly depends on the scale of its training data and the size of the model [1]. To create a better LLM you need to increase its size, and train it on more data. The architecture and configuration of an LLM can almost be neglected in comparison.

However, the intelligence of a model does not grow linearly with its size [1]. There is a diminishing return on increasing the size of LLMs. If you increase the size of model and training set by 10x, you will only get an increase in performance as the previous 10x increase accomplished. This explains why vast amounts of data are required to effectively train large language model.

ChatGPT 3 was trained on 500 billion tokens [2]. There are no official numbers (that I could find) on how much data was used to train ChatGPT 4, but rumors in the AI community state that it was trained on 13 trillion tokens. With this numbers the performance step from 3 to 4 took an 26x increase in data.

The estimation for the size of publicly available, de duplicated data is 320 trillion tokens [4]. This is 24.6x more data than ChatGPT4 was (likely) trained on. If these numbers are correct, the performance of LLMs will only increase as much as it increased from ChatGPT3 to ChatGPT4 before we ran out of data.

I doubt that this will be enough to reach AGI level intelligence.

Poisoned Data

Now you might say "we produce more data every day, models can get better in the future". We could just wait some decades, train new models from time to time and see a gradual increase in performance. And one day we suddenly have an AGI. But the release of ChatGPT created a problem. It poisoned any data collected after November 30, 2022.

The release of ChatGPT was the wet dream of every spammer, bot operator, troll and wannabe influencer. Suddenly you could create seemingly high quality content, indistinguishable from human written text, at virtually zero cost. Ever since the internet gets filled with LLM generated content. And this is a problem for all future trainings.

Models that are trained on their own generations (or data created by other models) start to forget, there performance declines [3]. Therefore you have to avoid training on AI generated content, otherwise the increase in data may decrease your performance. As it is virtually impossible to clean a dataset from AI generated text, any data collected after Nov. 2022 should be avoided. Maybe you can still use a few more months or years of data, but at some point more data will hurt the model more than it helps.

The publicly available data that we got now is all we will get to train an AGI. If we need more data it will have to collected at an exorbitant price to ensure it's not poisoned by AI generated data.

We reached a dead end

The size of current training sets and potentially available data shows that we will not reach AGI levels with the current state of the art approaches. Due to data poisoning we will not get substantially more usable data. Thereby AGI is (currently) not achievable. Opening pandora's box of accessible generative AI may killed our chance of creating an artificial general intelligence.

If we want to build an AGI we will have to do it with the data we have now.

#AI #AGI #LLM #GenAI


Szechuan Pfeffer Udon Nudeln

Servings: 1 Portion , Prep Time:

Ingredients

  • 100 g Udon Nudeln
  • Handvoll Sojaschnetzel
  • 1 kleine Zucchini
  • 1 kleine Karotte
  • 1 EL Szechuan Pfeffer
  • 1 Knoblauchzehe
  • 1 cm Ingwer
  • 1 EL schwarze Bohnenpaste
  • 1 EL Sojasauce
  • 1 EL Agavendicksaft
  • 200 ml Wasser

Instructions

Dieses Rezept war ein bisschen frei Schnauze, ist aber wirklich gut geworden. Die Mengen sind im nachhinein geschätzt, also abschmecken und gegebenfalls anpassen.

Szechuan Pfeffer Udon Nudeln

  • Sojaschnetzeln kochen
  • Knoblauch und Ingwer fein hacken
  • Alle Gewürze und Wasser mischen
  • Karotten und Zucchini in dünne Scheiben schneiden
  • Nudeln nach Packung kochen
  • Sojaschnetzel anbraten
  • Zucchini und Karotten für 2 Minuten mitanbraten
  • Mit Gewürzmischung ablöschen
  • Sauce reduzieren und Nudeln unterheben

#vegan #nudeln #udon #zucchini #karotten #szechuan #szechuanpfeffer


Random Tables for DnD 5e

List of random tables I've created to quickly create shops or treasures when DMing.

Spells

Magic Items

Armor

Rings

Rods

Staves

Weapons

Wondrous Items

Generators

Gamemaster Guide Tables

#DnD #5e #random #generator #tabletop


Nudeln mit Erbsen

Servings: 3 Portionen , Prep Time:

Ingredients

  • 250g Nudeln
  • 200g TK Erbsen
  • 200ml vegane Sahne
  • 1 Zwiebel
  • 1 Knoblauchzehe
  • 50g Hefeflocken
  • 1 EL Olivenöl
  • 1 EL Tomatenmark
  • 1 TL Paprikapulver
  • 1 TL Oregano
  • Salz und Pfeffer

Instructions

Ein einfaches und schnelles Nudelgericht ideal für Kleinkinder. Schmeckt gut mit veganem Parmesan.

  • Zwiebel würfeln und Knoblauch fein hacken
  • Nudeln kochen
  • Zwiebeln und Knoblauch in Olivenöl glasig anbraten
  • Tomatenmark für 2-3 Minuten mitanbraten
  • Erbsen, Hefeflocken und Gewürze hinzugeben, kurz mit anbraten
  • Sahne und 100ml Wasser hinzugeben und gut verrühren
  • Sauce unter gelegentlichem Rühren aufkochen
  • Nudeln abgießen und alles gut durchmischen
  • Mit Salz und Pfeffer abschmecken

#vegan #rezept #nudeln #erbsen


Veganer Pesto Nudelsalat mit Rucola

Servings: 4 Portionen , Prep Time:

Ingredients

  • 500g Fusilli Nudeln
  • 250g Cherry Tomaten
  • 125g Rucola
  • 1 Glas veganes grünes Pesto
  • 2 EL Agavendicksaft
  • 2 EL Senf, mittelscharf
  • 50g veganer Parmesan (optional)

Instructions

Die vegane Version meines Lieblings-Nudelsalats.

  • Nudeln kochen und abkühlen lassen oder mit kalten Wasser abkühlen
  • Rucola waschen und abtropfen lassen
  • Tomaten halbieren oder vierteln
  • Nudeln, Tomaten, Pesto, Agavendicksaft und Senf vermischen
  • Mit veganem Parmesan abschmecken (optional)
  • Rucola untermischen

#vegan #rezept #nudelsalat #einfach


Microformat Categories for Hashtags in owl-blogs

I recently added hashtag support to owl-blogs. The initial reason for this was to to make post more discoverable via ActivityPub, but I found it helpful to further categories posts. The implementation was quiet simple. I use the markdown renderer goldmark which has a plugin for hashtags.

As tags are also part of microformats2, I wanted to mark the hashtags accordingly. This is currently not possible with the hashtag extension.

I've extended this to allow adding arbitrary attributes to the link tag (Related Pull Request). Until this is merged into the main repository I'll use my own version, which can be done by adding a replace directive to the go.mod

replace go.abhg.dev/goldmark/hashtag => github.com/H4kor/goldmark-hashtag v0.0.0-20240619193802-bec327f5be38

#go #dev #owlblogs #markdown #indieweb




Thumbnails for owl-blogs

I'm currently working on the thumbnail implementation for #owl-blogs and deployed the feature branch to my blog.

Thumbnails use the same file format as their parent files assuming the user already chose the best format for their images. The thumbnails are created with a width of 620px, equal to the content width of the blogs main body. If the image is already small enough the image data is simply copied.

The URL of a thumbnail can simply be generated by replacing /media/ with /thumbnail/.

I will still write some tests and see if any errors occur on my blog before merging this feature into the main branch.

#dev #blog


Serving Binary Files from SQLite

My blog software (owl-blogs) uses a single SQLite database to store everything, including all files uploaded. I'm aware that storing large files in a relational database isn't best practice. It started out as a placeholder implementation, but I liked the idea to have a single file I can backup.

One reason against storing binary blobs in relational databases often stated is read performance, but I didn't find any benchmarks supporting this claim. Therefore I built a small test setup to see the difference between serving binary files out of a SQLite database vs serving from the file system directly.

As my blog is written in Go, I created the a simple server similar to my blog. It uses sqlx and go-sqlite3 for the database handling and net/http for the static file server

package main

import (
	"log"
	"net/http"

	"github.com/jmoiron/sqlx"
	_ "github.com/mattn/go-sqlite3"
)

type sqlBinaryFile struct {
	Data []byte `db:"data"`
}

type sqlHandler struct {
	Db *sqlx.DB
}

func (h *sqlHandler) ServeHTTP(w http.ResponseWriter, r *http.Request) {
	id := r.PathValue("filename")
	var sqlFile sqlBinaryFile
	h.Db.Get(&sqlFile, "SELECT data FROM files WHERE id = ?", id)
	w.Write(sqlFile.Data)
}

func main() {

	db := sqlx.MustOpen("sqlite3", "files.db")
	sql := &sqlHandler{Db: db}

	fs := http.StripPrefix("/dir/", http.FileServer(http.Dir("./static")))
	http.Handle("/dir/", fs)
	http.Handle("/sqlite/{filename}", sql)

	log.Print("Listening on :3000...")
	err := http.ListenAndServe(":3000", nil)
	if err != nil {
		log.Fatal(err)
	}
}

As a test set I created 2000 files between 200kb and 4MB in size using a simple python script:

import os
import random

for i in range(2000):
    os.system(f"head -c {(random.randint(200, 4000))}K </dev/urandom > static/{i:05d}.bin")

The SQLite database was created with this script:

import os
import sqlite3

os.remove("files.db")
con = sqlite3.connect("files.db")
cur = con.cursor()

cur.execute("CREATE TABLE files( id VARCHAR(255) PRIMARY KEY, data BLOB NOT NULL )")

for f in os.listdir("static"):
    print(f)
    data = open("static/" + f, "rb").read()
    cur.execute("INSERT INTO files(id, data) VALUES (?, ?)", (f, data))

con.commit()

To benchmark the server I created two files listing all file URLs (one for sqlite, one fot filesystem) and used siege to run the benchmark with this configuration.

siege -f urls_sqlite.txt -c 1 -b --time=10s -j

The test was executed on my laptop:

  • CPU: Intel(R) Core(TM) i7-8565U CPU @ 1.80GHz
  • CPU max MHz: 4600,0000
  • Memory: 16 GB

I ran the test with different concurrency and plotted the results:

Two plots comparing transactions and response time between SQLite and filesystem

For a low throughput system (such as my blog) the difference between SQLite and the filesystem is small enough to not care about. The possible throughput (transaction/second) of the filesystem is ~2.3 times higher. The response time grows slower with increased concurrency.

For the time being I will stick with my SQLite solution. Once my blog gets really popular I can easily change the implementation of the binary repository.

#hosting #sqlite #go #benchmark #server