How to release memory from a record's field?

Hello, hope you guys are doing great. I was wondering how can I release the memory used by a record’s field. For instance, in my server, I’m holding a file’s content like:

#state{file_content = <<...>>}

And when I no longer need it, I’m doing:

#state{file_content = <<>>}

Expecting that it will free/release the memory but I’m wondering if this actually works and/or if it’s the correct approach.

Thanks!

That will work but you won’t see the effect until the process has been garbage collected.

Hey @mikpe thanks for the reply. With that in mind, I wonder if it would made sense to extract the file binary into its own gen server. That would also help with logs when things crash. Right now, if it crashes the logs are terrible to look at (you get the entire binary printed)

Perhaps something simple like:

-record(state, {file_content = <<>>}).

% init
% FileContent = read_file(...)

handle_call({get_chunk, OffsetInBytes, ChunkSizeInBytes}, FromPID, State) ->
    Chunk = read_file_chunk(State#state.file_content, OffsetInBytes, ChunkSizeInBytes),
    {reply, Chunk, State}.

handle_cast({stop}, State) ->
    {stop, normal, State}.

to get rid of unwanted output in crashes/logs: use the format_status/1 callback gen_server — OTP 29.1.1 (stdlib 8.1)

A few questions…

Is this an on-disk file that you’re loading into memory on init?

What do the access patterns look like?

Why not just keep it on disk and use the file module functions to read on demand?

If it has to be ‘in process’, could you not store it in a persistent_term rather than state? That way it won’t throw up in your logs and would survive other process crashes. Unfortunately does not lend itself well to chunked reading. Storing the file data in ETS would be similar issue

Thank you @sg2342 that’s quite handy.

Hello @LeonardB ,

Is this an on-disk file that you’re loading into memory on init?

Yes, its a file I’m loading from disk.

What do the access patterns look like?

I’m running a game server emulator where players can create their own lobbies and you tell which level/scenario players will play. In a lobby you can have up to 12 players. When players join, the game downloads (if needed) that level/scenario file in chunks. When players are ready, you can start the game and at this point, I no longer need to have the file in memory which is why I was looking to reclaim that memory.

Why not just keep it on disk and use the file module functions to read on demand?

Initially, I was using pread and reading/processing chunks on demand but the CPU went to 100% quite easily. I don’t recall if it affected my users though but probably.

If it has to be ‘in process’, could you not store it in a persistent_term rather than state? That way it won’t throw up in your logs and would survive other process crashes. Unfortunately does not lend itself well to chunked reading. Storing the file data in ETS would be similar issue

Honestly I would like to keep things super simple and just use gen_servers. I think having a separate gen_server like I posted would work fine. I will test it and keep you guys posted.

Thanks again!

Avoid holding the entire file content in memory, the best approach is to “not read it all at once in the first place”. You can change your logic to read chunks on demand from the file system.

% Instead of reading the whole file into state, store the file path.
-record(state, {file_path :: string()}).

% Later, when a chunk is needed:
handle_call({get_chunk, Offset, Size}, _From, State) ->
    {ok, Fd} = file:open(State#state.file_path, [read, raw, binary]),
    {ok, _} = file:position(Fd, Offset),
    {ok, Chunk} = file:read(Fd, Size),
    file:close(Fd),
    {reply, Chunk, State}.

This approach keeps your process heap (and the node’s total memory usage) minimal, as you only ever hold the chunk you are currently processing.

Thanks @keyhan . I originally had it this way but the CPU spiked to 100% quite easily (now that I think about it, I think I was opening/reading/closing for every chunk). One note though, I’m using file:pread since according to the documentation its faster than manually calling position + read.

Combines position/2 and read/2 in one operation, which is more efficient than calling them one at a time.

In any case, I went back to the drawing board and decided not to load the entire file in RAM. The reason is that I’m running on a very modest server (1vcore, 2GB RAM) and each level/scenario can weight between 1MB all the way to 100MB which would limit how many lobbies I can create quite drastically.

What I ended up doing is, I’m holding the file’s FD in the state and processing the chunks on demand. When I no longer need it, I close the file.

Having the FD, means I don’t have to open/close every time I need to process a chunk.

When I need to send the chunks to a player, I have to split and send many chunks in sizes up to 1KB (I can’t do bigger because it crashes the client) For this, I read ChunksToSend * 1024 from the file, and from that resulting binary I split and send the processed chunks (as opposed to reading 1024, process, repeat).

Will keep you posted. Hopefully it works better now. Thanks everybody!

EDIT: went back and check for this:

I originally had it this way but the CPU spiked to 100% quite easily (now that I think about it, I think I was opening/reading/closing for every chunk)

And indeed, I was opening, reading up to 1024 bytes, process, close for every chunk so yeah, maybe that was the cause of the CPU spikes haha.