<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Artem Golubin</title><link>https://rushter.com/blog/feed/</link><description>Python, Machine learning, NLP, websec, etc.</description><language>en</language><lastBuildDate>Sun, 20 Sep 2026 18:43:08 -0000</lastBuildDate><item><title>Real attackers don't think in bug bounty scopes</title><link>https://rushter.com/blog/security-bb-and-scopes/</link><description>&lt;p&gt;A recent post on &lt;a href="https://www.hacktron.ai/blog/hacking-openai" target="_blank"&gt;Hacking OpenAI&lt;/a&gt; and a bug bounty of $6,500 reminded me
of an old problem. A security company was able to hack into private OpenAI repositories, but only got a small
bounty. A company with a valuation of $852 billion paid a few pennies to researchers. In comparison, such data
could be sold for hundreds of thousands on the black market.&lt;/p&gt;
&lt;p&gt;For those who don't know, a bug bounty program is a way for companies to pay white-hat hackers (security researchers)
to find vulnerabilities in their services and software.
Bug bounties allow researchers to perform security research in a legal way and get paid for it.
The hack was initiated through a hosted Discourse forum, but that forum was out of the scope of OpenAI's bug bounty program.&lt;/p&gt;
&lt;p&gt;In bug bounty programs, there is usually a list of services the company wants to be tested and the rest of the services
do not receive any bounty. In the case of the OpenAI hack, accessing the private data (from repositories) was in scope,[......]</description><pubDate>Sun, 20 Sep 2026 18:43:08 -0000</pubDate><guid>https://rushter.com/blog/security-bb-and-scopes/</guid></item><item><title>Astroturfing on Reddit and the dead internet</title><link>https://rushter.com/blog/astroturfing/</link><description>&lt;p&gt;Reddit used to be a good source to find real recommendations. Well, not anymore.
Years ago, I used to append "reddit" to my search queries to find real discussions about products.&lt;/p&gt;
&lt;p&gt;If you search for managed Postgres on Reddit, you may stumble on posts like this:&lt;/p&gt;
&lt;p&gt;&lt;img src="/static/uploads/img/2026/astro/reddit1.png" class="ui centered image" &gt;&lt;/p&gt;

&lt;p&gt;If you look at the profile of the person who created it, their posting history is usually private:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;u/xxx likes to keep their posts hidden, but check out their stats to learn more about them.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;But when you do some research and look at their other posts from Reddit dumps to find their real identity, you see this:&lt;/p&gt;
&lt;p&gt;&lt;img src="/static/uploads/img/2026/astro/linkedin1.png" class="ui centered image" &gt;&lt;/p&gt;

&lt;p&gt;By diving deeper, you notice that there are around 10 accounts that constantly ask about managed Postgres or recommend
a specific provider. They are all connected to the same company.[......]</description><pubDate>Fri, 11 Sep 2026 19:02:08 -0000</pubDate><guid>https://rushter.com/blog/astroturfing/</guid></item><item><title>Writing arenas in Rust from scratch</title><link>https://rushter.com/blog/rust-memory-arenas/</link><description>&lt;p&gt;Custom memory allocations and arenas are notoriously hard to deal with in Rust due to the ownership model.&lt;/p&gt;
&lt;p&gt;In a lot of projects, they can significantly improve performance.
For example, let's take databases.
Sending a query to a database produces a lot of temporary objects, and if the default allocator is used,
you will need to allocate and deallocate thousands of small objects for each query, which is wasteful.&lt;/p&gt;
&lt;p&gt;Almost every database engine uses custom memory pools to improve performance.
Some go so far as to &lt;a href="https://github.com/tigerbeetle/tigerbeetle/blob/main/docs/ARCHITECTURE.md#static-memory-allocation" target="_blank"&gt;statically allocate&lt;/a&gt;
all the memory before the program starts.&lt;/p&gt;
&lt;p&gt;Another good example is scripting languages. I have &lt;a href="/blog/python-memory-managment/" target="_blank"&gt;an article&lt;/a&gt;
 on how Python uses arenas to reduce the number of allocations for Python types. A simple web app in Python can
&lt;a href="/blog/python-object-allocation-statistics/" target="_blank"&gt;allocate&lt;/a&gt; more than a &lt;strong&gt;million&lt;/strong&gt; of objects after 100 HTTP requests.[......]</description><pubDate>Mon, 27 Jul 2026 17:02:08 -0000</pubDate><guid>https://rushter.com/blog/rust-memory-arenas/</guid></item><item><title>Domain registration information should be transparent</title><link>https://rushter.com/blog/domain-whois-information/</link><description>&lt;p&gt;The single best signal for identifying a phishing domain is the domain registration date, and yet
domain registrars still limit access to such data.&lt;/p&gt;
&lt;p&gt;For certificates, we have &lt;a href="https://certificate.transparency.dev/" target="_blank"&gt;Certificate Transparency Logs&lt;/a&gt; where
every issued TLS certificate is publicly visible within seconds of issuance.&lt;/p&gt;
&lt;p&gt;For domains, we used to use WHOIS protocol, which returned plaintext data:&lt;/p&gt;
&lt;p&gt;&lt;img src="/static/uploads/img/2026/whois.png" class="ui centered image" &gt;&lt;/p&gt;

&lt;p&gt;It was pretty hard to deal with, because every domain zone used its own format that was hard to parse.&lt;/p&gt;
&lt;p&gt;In the last five years,
a new &lt;a href="https://en.wikipedia.org/wiki/Registration_Data_Access_Protocol" target="_blank"&gt;RDAP&lt;/a&gt; protocol has grown in popularity.
Instead of plaintext, it now outputs JSON data.&lt;/p&gt;
&lt;p&gt;It became a mandatory protocol in 2025 and more than 1440 TLDs (domain zones) support it now.
Despite the new protocol, the problem persists - registrars are still rate limiting access.&lt;/p&gt;[......]</description><pubDate>Wed, 22 Jul 2026 15:14:08 -0000</pubDate><guid>https://rushter.com/blog/domain-whois-information/</guid></item><item><title>Hexora v0.3: New features and improvements</title><link>https://rushter.com/blog/hexora-update/</link><description>&lt;p&gt;Recently, I've improved my Python library, &lt;a href="https://github.com/rushter/hexora" target="_blank"&gt;hexora&lt;/a&gt;.
I wrote it to detect &lt;a href="/blog/python-code-exec/" target="_blank"&gt;malicious Python code&lt;/a&gt; using static analysis.&lt;/p&gt;
&lt;p&gt;In the new v.0.3.0 release, I've added new detections, and we now also use a simple machine learning model to analyze the whole file.
The machine learning model uses code structure features, semantic features, and static code analysis to assess the entire Python file.&lt;/p&gt;
&lt;p&gt;Although the model can detect malicious code without any detections coming from static analysis,
its main use case is to filter false positives.&lt;/p&gt;
&lt;p&gt;I've been testing it against newly published PyPI packages and it detects 2-10 new malicious packages each day.&lt;/p&gt;
&lt;p&gt;&lt;img src="/static/uploads/img/2026/pypi_security.png" class="ui centered image" &gt;&lt;/p&gt;

&lt;p&gt;Due to the number of published packages, before the machine learning model, I was getting around 5-10 false positives for 1[......]</description><pubDate>Thu, 25 Jun 2026 15:14:08 -0000</pubDate><guid>https://rushter.com/blog/hexora-update/</guid></item><item><title>Using local ClickHouse for data processing</title><link>https://rushter.com/blog/clickhouse-data-processing/</link><description>&lt;p&gt;I did a lot of data engineering work in my career.&lt;/p&gt;
&lt;p&gt;When you work a lot with data, you often get quick requests to extract some cold data and process it.
Since the data is cold, it usually resides on S3.&lt;/p&gt;
&lt;p&gt;For example, one of the typical requests in the past was to count unique values from an old MongoDB backup or CSV dump with 100GB of compressed data.&lt;/p&gt;
&lt;p&gt;I quickly learned that using throwaway Python scripts would not work well.
Often, the data is too big to fit into memory to sort and deduplicate.
Maintaining a Spark cluster was not worth it for this kind of work.&lt;/p&gt;
&lt;p&gt;So, what I would do is something like this:&lt;/p&gt;
&lt;div class="codehilite"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;&lt;span class="nv"&gt;LC_ALL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;C&lt;span class="w"&gt; &lt;/span&gt;aws&lt;span class="w"&gt; &lt;/span&gt;s3&lt;span class="w"&gt; &lt;/span&gt;cp&lt;span class="w"&gt; &lt;/span&gt;s3://bucket/data.json.gz&lt;span class="w"&gt; &lt;/span&gt;-&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="se"&gt;\&lt;/span&gt;
&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="p"&gt;|&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;gzip&lt;span class="w"&gt; &lt;/span&gt;-dc&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="se"&gt;\&lt;/span&gt;[......]</description><pubDate>Sat, 06 Jun 2026 14:14:08 -0000</pubDate><guid>https://rushter.com/blog/clickhouse-data-processing/</guid></item><item><title>NULLs in ClickHouse can hurt performance</title><link>https://rushter.com/blog/clickhouse-nulls/</link><description>&lt;p&gt;When coming from relational databases, NULLs are the go-to for optional fields.
Using them in ClickHouse can lead to unexpected and often unnoticeable performance degradation.
This article explain why.&lt;/p&gt;
&lt;h4&gt;PostgreSQL&lt;/h4&gt;
&lt;p&gt;When using null values in PostgreSQL, you rarely notice any difference.
In PG, columns are nullable by default and you can index them.&lt;/p&gt;
&lt;p&gt;Internally, each row in PostgreSQL has a bitmap that indicates which columns are NULL. It's only present when
there are null values in a particular row using a bit flag.&lt;/p&gt;
&lt;p&gt;PostgreSQL is a row-oriented database, so when you read a row, you read all the columns together.&lt;/p&gt;
&lt;h4&gt;ClickHouse&lt;/h4&gt;
&lt;p&gt;Unlike PostgreSQL, ClickHouse is a columnar database.
Instead of storing data by rows, it organizes them by columns.
Each column is stored separately as a contiguous block of data.&lt;/p&gt;
&lt;p&gt;Let's suppose we have a table that stores HTTP logs.
We want to store the visitor's user ID, which can be empty for anonymous users.&lt;/p&gt;[......]</description><pubDate>Wed, 03 Jun 2026 14:14:08 -0000</pubDate><guid>https://rushter.com/blog/clickhouse-nulls/</guid></item><item><title>PyPI packages are increasing rapidly</title><link>https://rushter.com/blog/pypi-packages/</link><description>&lt;p&gt;PyPI is the main repository for Python packages.
One thing that I've noticed recently is the number of published packages per week.&lt;/p&gt;
&lt;p&gt;Let's look at published counts of new package versions per week:&lt;/p&gt;
&lt;p&gt;&lt;img src="/static/uploads/img/2026/pypi_stats_weekly.png" class="ui centered image" &gt;&lt;/p&gt;

&lt;p&gt;There are some dips in the data, but that's because of how the &lt;a href="https://clickpy.clickhouse.com/" target="_blank"&gt;data&lt;/a&gt; was collected.
We can see a clear increase in the number of published packages, especially in the last few months.&lt;/p&gt;
&lt;p&gt;Because of AI, the number of packages published per week has increased by 30% since 2025.&lt;/p&gt;
&lt;p&gt;I'm working on &lt;a href="https://github.com/rushter/hexora" target="_blank"&gt;hexora&lt;/a&gt;, a library that detects malicious Python code in packages.
It monitors newly published PyPI packages in real time and analyzes them.&lt;/p&gt;
&lt;p&gt;A lot of packages, that have been published recently, are purely vibecoded, and they trigger false positive detections when my tool analyzes them.[......]</description><pubDate>Sun, 17 May 2026 14:14:08 -0000</pubDate><guid>https://rushter.com/blog/pypi-packages/</guid></item><item><title>The rise of malicious repositories on GitHub</title><link>https://rushter.com/blog/github-malware/</link><description>&lt;p&gt;There is an ongoing surge of malicious repositories on GitHub, and the sad thing about it is that
GitHub seems not to care much.&lt;/p&gt;
&lt;p&gt;About 10 days ago, I searched for a repo on DuckDuckGo and stumbled upon a fake GitHub repo.
It mimics a legitimate repository, but instead of providing usual releases, it only provides malicious Windows binaries.
Linux/MacOS binaries are not available, and the information on how to build the project was removed from the README file.&lt;/p&gt;
&lt;p&gt;The description was also altered using LLMs, removing a lot of technical details.&lt;/p&gt;
&lt;p&gt;I reported this repository to GitHub, explaining the problem and showing the report from VirusTotal.
To this day, the repository is still there, and the binaries are still available for download.&lt;/p&gt;
&lt;p&gt;The repo has been active for two months. The README gets constantly updated every hour so that it will appear in the
GitHub search higher.&lt;/p&gt;
&lt;p&gt;&lt;img src="/static/uploads/img/2026/vt.png"&gt;&lt;/p&gt;
&lt;p&gt;Today, I saw another case of this on &lt;a href="https://x.com/rebane2001/status/2033208600072425780" target="_blank"&gt;X&lt;/a&gt;,[......]</description><pubDate>Sun, 15 Mar 2026 17:14:08 -0000</pubDate><guid>https://rushter.com/blog/github-malware/</guid></item><item><title>Do not fall for complex technology</title><link>https://rushter.com/blog/complex-tech/</link><description>&lt;p&gt;Fifteen years ago, I wanted to set up a note-taking system.
At the time, Evernote was the tool everyone was talking about, so choosing it seemed like the right and easy decision.&lt;/p&gt;
&lt;p&gt;After storing around 500 notes for eight years, Evernote became a mess to use.
It was bloated, heavily monetized, and slow to work with. So I wanted to switch.&lt;/p&gt;
&lt;p&gt;About that time came Notion. Everyone was talking about it. I jumped on the bandwagon and migrated a few hundred of
my notes to it that were still relevant. It did not even occur to me that switching from a bloated and slow app to
a web app would result in a similar outcome later. I followed a popular choice again.&lt;/p&gt;
&lt;p&gt;After struggling for a year, I switched to Markdown notes and a plugin for an editor that renders inline images.
I'm still using this to this day. It is simple, and I will be able to open my notes 10-20 years later.
I can edit them in any editor. It works offline and does not depend on commercial products.
I encrypt my notes locally so they can be stored safely on any cloud service.&lt;/p&gt;[......]</description><pubDate>Thu, 22 Jan 2026 17:14:08 -0000</pubDate><guid>https://rushter.com/blog/complex-tech/</guid></item></channel></rss>