restic

Commit Graph

Author	SHA1	Message	Date
greatroar	d129baba7a	repository: Reuse buffers in Repository.LoadUnpacked This method had a buffer argument, but that was nil at all call sites. That's removed, and instead LoadUnpacked now reuses whatever it allocates inside its retry loop.	2023-01-30 22:01:01 +01:00
Michael Eischer	1adf28a2b5	repository: properly return invalid data error in LoadUnpacked The retry backend does not return the original error, if its execution is interrupted by canceling the context. Thus, we have to manually ensure that the invalid data error gets returned. Additionally, use the retry backend for some of the repository tests, as this is the configuration which will be used by restic.	2023-01-14 17:57:02 +01:00
Michael Eischer	6d9675c323	repository: cleanup error message on invalid data The retry printed the filename twice: ``` Load(<lock/04804cba82>, 0, 0) returned error, retrying after 720.254544ms: load(<lock/04804cba82>): invalid data returned ``` now the warning has changed to ``` Load(<lock/04804cba82>, 0, 0) returned error, retrying after 720.254544ms: invalid data returned ```	2023-01-14 17:57:02 +01:00
Michael Eischer	90fb6f70b4	Merge pull request #4089 from greatroar/errors Clean up error handling further	2022-12-24 10:41:56 +01:00
greatroar	b150dd0235	all: Replace some errors.Wrap calls by errors.WithStack Mostly changed the ones that repeat the name of a system call, which is already contained in os.PathError.Op. internal/fs.Reader had to be changed to actually return such errors.	2022-12-17 09:41:07 +01:00
greatroar	c0b5ec55ab	repository: Remove empty cleanup functions in tests TestRepository and its variants always returned no-op cleanup functions. If they ever do need to do cleanup, using testing.T.Cleanup is easier than passing these functions around.	2022-12-11 11:06:25 +01:00
Michael Eischer	40ac678252	backend: remove Test method The Test method was only used in exactly one place, namely when trying to create a new repository it was used to check whether a config file already exists. Use a combination of Stat() and IsNotExist() instead.	2022-12-03 11:28:10 +01:00
Michael Eischer	ff7ef5007e	Replace most usages of ioutil with the underlying function The ioutil functions are deprecated since Go 1.17 and only wrap another library function. Thus directly call the underlying function. This commit only mechanically replaces the function calls.	2022-12-02 19:36:43 +01:00
Michael Eischer	a1eb923876	remove no longer necessary conditional compiles	2022-11-27 13:18:44 +01:00
Alexander Neumann	8dd95b710e	Merge pull request #3992 from MichaelEischer/err-on-invalid-compression Return error if RESTIC_COMPRESSION env variable is invalid	2022-11-04 19:41:34 +01:00
greatroar	137f0bc944	repository: Fix benchmarkSaveAndEncrypt	2022-10-29 23:09:17 +02:00
Michael Eischer	01f0db4e56	return error if RESTIC_COMPRESSION env variable is invalid	2022-10-29 22:03:39 +02:00
Michael Eischer	c4fc5c97f9	prune: Use a single CountedBlobSet to track blobs The set covers necessary, existing and duplicate blobs. This removes the duplicate sets used to track whether all necessary blobs also exist. This reduces the memory usage of prune by about 20-30%.	2022-10-22 18:45:12 +02:00
Michael Eischer	8d62a7adb4	identify keys by ID and not name	2022-10-15 16:07:43 +02:00
Michael Eischer	02634dce7a	restic: change Find to return ids That way consumers no longer have to manually convert the returned name to an id.	2022-10-15 16:06:54 +02:00
Michael Eischer	2e3f1c08c5	repository: split index into a separate package	2022-10-08 21:15:34 +02:00
Michael Eischer	5760ba6989	Merge pull request #3949 from MichaelEischer/simplify-mixedpacks repository: remove IsMixedPack and add replacement for checker	2022-10-08 21:14:14 +02:00
Michael Eischer	4bb5240720	repository: remove unused PrefixLength	2022-10-03 12:15:53 +02:00
Michael Eischer	999fe29976	repository: hide prepareCache	2022-10-03 12:15:53 +02:00
Michael Eischer	ddcf549eba	repository: remove IsMixedPack and add replacement for checker Repositories with mixed packs are probably quite rare by now. When loading data blobs from a mixed pack file, this will no longer trigger caching that file. However, usually tree blobs are accessed first such that this shouldn't make much of a difference. The checker gets a simpler replacement.	2022-10-03 12:03:59 +02:00
Michael Eischer	5c6b6edefe	retry index, lock and snapshot loading on hash mismatch	2022-09-25 11:35:35 +02:00
Michael Eischer	78d2312ee9	Merge pull request #3854 from MichaelEischer/sparsefiles restore: Add support for sparse files	2022-09-24 22:04:02 +02:00
Michael Eischer	c147422ba5	repository: special case SaveBlob for all zero chunks Sparse files contain large regions containing only zero bytes. Checking that a blob only contains zeros is possible with over 100GB/s for modern x86 CPUs. Calculating sha256 hashes is only possible with 500MB/s (or 2GB/s using hardware acceleration). Thus we can speed up the hash calculation for all zero blobs (which always have length chunker.MinSize) by checking for zero bytes and then using the precomputed hash. The all zeros check is only performed for blobs with the minimal chunk size, and thus should add no overhead most of the time. For chunks which are not all zero but have the minimal chunks size, the overhead will be below 2% based on the above performance numbers. This allows reading sparse sections of files as fast as the kernel can return data to us. On my system using BTRFS this resulted in about 4GB/s.	2022-09-24 21:39:39 +02:00
Michael Eischer	1ebd57247a	repository: optimize MasterIndex.Each Sending data through a channel at very high frequency is extremely inefficient. Thus use simple callbacks instead of channels. > name old time/op new time/op delta > MasterIndexEach-16 6.68s ±24% 0.96s ± 2% -85.64% (p=0.008 n=5+5)	2022-09-24 12:21:59 +02:00
Michael Eischer	825b95e313	repository: add benchmark for MasterIndex.Each	2022-09-24 12:21:59 +02:00
Michael Eischer	7682149c9d	repository: cleanup copy connection count check	2022-08-28 11:40:56 +02:00
Michael Eischer	b03277ead5	repository: don't hang when copying using a single connection	2022-08-28 11:40:31 +02:00
MichaelEischer	bee15dd555	Merge pull request #3879 from MichaelEischer/mem-optimize Some random (minor) memory-allocation optimizations	2022-08-26 20:33:02 +02:00
Michael Eischer	cc4728d287	repository: Do not report ignored packs in EachByPack Ignored packs were reported as an empty pack by EachByPack. The most immediate effect of this is that the progress bar for rebuilding the index reports processing more packs than actually exist.	2022-08-21 10:38:40 +02:00
Michael Eischer	7a992fc794	repository: Reduce buffer reallocations in ForAllIndexes Previously the buffer was grown incrementally inside `repo.LoadUnpacked`. But we can do better as we already know how large the index will be. Allocate a bit more memory to increase the chance that the buffer can be reused in the future.	2022-08-19 21:13:40 +02:00
Michael Eischer	77b1980d8e	repository: MasterIndex.Packs: reduce allocations	2022-08-19 21:10:43 +02:00
Michael Eischer	6ff9517e45	repository: MasterIndex.ListPacks / Index.EachByPack allow earlier GC Allow earlier garbage collection of some of the intermediate data structures.	2022-08-19 21:06:33 +02:00
Michael Eischer	f414db987d	gofmt all files Apparently the rules for comment formatting have changed with go 1.19.	2022-08-19 19:12:26 +02:00
Michael Eischer	7266f07c87	repository: StreamPack in parts if there are too large gaps For large pack sizes we might be only interested in the first and last blob of a pack file. Thus stream a pack file in multiple parts if the gaps between requested blobs grow too large.	2022-08-05 23:48:36 +02:00
Michael Eischer	1b076cda97	rename option to --pack-size	2022-08-05 23:47:43 +02:00
Kyle Brennan	1e3f05c3f1	repository: prevent header overfill	2022-08-05 23:47:12 +02:00
Michael Eischer	0a6fa602c8	add option for setting min pack size	2022-08-05 23:47:12 +02:00
Michael Eischer	73053674d9	repository: Test fallback to existing blobs	2022-07-30 17:37:07 +02:00
Michael Eischer	623770eebb	repository: try to recover from invalid blob while repacking If a blob that should be kept is invalid, Repack will now try to request the blob using LoadBlob. Only return an error if that fails.	2022-07-30 17:37:07 +02:00
MichaelEischer	443cc49afd	Merge pull request #3830 from MichaelEischer/cleanup-repo Extract Load/SaveTree/JSONUnpacked from repository	2022-07-23 10:46:13 +02:00
Michael Eischer	9729e6d7ef	backend: extract readerat from restic package	2022-07-17 15:29:09 +02:00
Michael Eischer	8c11fc3ec9	crypto: move crypto buffer helpers	2022-07-17 13:42:23 +02:00
Michael Eischer	89d3ce852b	repository: extract Load/StoreJSONUnpacked A Load/Store method for each data type is much clearer. As a result the repository no longer needs a method to load / store json.	2022-07-17 13:22:00 +02:00
Michael Eischer	fbcbd5318c	repository: extract LoadTree/SaveTree The repository has no real idea what a Tree is. So these methods never belonged there.	2022-07-17 13:11:28 +02:00
Lorenz Bausch	d6e3c7f28e	Wording: change repo to repository	2022-07-08 20:05:35 +02:00
Michael Eischer	6f53ecc1ae	adapt workers based on whether an operation is CPU or IO-bound Use runtime.GOMAXPROCS(0) as worker count for CPU-bound tasks, repo.Connections() for IO-bound task and a combination if a task can be both. Streaming packs is treated as IO-bound as adding more worker cannot provide a speedup. Typical IO-bound tasks are download / uploading / deleting files. Decoding / Encoding / Verifying are usually CPU-bound. Several tasks are a combination of both, e.g. for combined download and decode functions. In the latter case add both limits together. As the backends have their own concurrency limits restic still won't download more than repo.Connections() files in parallel, but the additional workers can decode already downloaded data in parallel.	2022-07-03 12:19:26 +02:00
Michael Eischer	753e56ee29	repository: Limit to a single pending pack file Use only a single not completed pack file to keep the number of open and active pack files low. The main change here is to defer hashing the pack file to the upload step. This prevents the pack assembly step to become a bottleneck as the only task is now to write data to the temporary pack file. The tests are cleaned up to no longer reimplement packer manager functions.	2022-07-02 22:42:34 +02:00
Michael Eischer	120ccc8754	repository: Rework blob saving to use an async pack uploader Previously, SaveAndEncrypt would assemble blobs into packs and either return immediately if the pack is not yet full or upload the pack file otherwise. The upload will block the current goroutine until it finishes. Now, the upload is done using separate goroutines. This requires changes to the error handling. As uploads are no longer tied to a SaveAndEncrypt call, failed uploads are signaled using an errgroup. To count the uploaded amount of data, the pack header overhead is no longer returned by `packer.Finalize` but rather by `packer.HeaderOverhead`. This helper method is necessary to continue returning the pack header overhead directly to the responsible call to `repository.SaveBlob`. Without the method this would not be possible, as packs are finalized asynchronously.	2022-07-02 22:42:34 +02:00
Michael Eischer	04c23fa95d	rebuild-index: correctly rebuild index for mixed packs For mixed packs, data and tree blobs were stored in separate index entries. This results in warning from the check command and maybe other problems.	2022-07-02 19:24:02 +02:00
Michael Eischer	a6e9e08034	Account for pack header overhead at each entry This will miss the pack header crypto overhead and the length field, which only amount to a few bytes per pack file.	2022-07-02 18:55:58 +02:00

1 2 3 4 5 ...

290 Commits